Baidu has released vLLM Kunlun, a community-maintained plugin that enables the vLLM inference engine to run on Kunlun XPU hardware. This integration allows for seamless execution of various open-source LLMs, including Transformer-based, Mixture-of-Expert (MoE), embedding, and multimodal models, on Kunlun hardware. The plugin supports a wide range of models and features such as quantization, LoRA fine-tuning, optimized attention mechanisms, speculative decoding, and tensor parallelism, while also offering an OpenAI-compatible API for model serving. AI
IMPACT Enables broader adoption of LLMs on specialized hardware, potentially improving inference efficiency and cost-effectiveness.
RANK_REASON This is a hardware plugin for an existing inference engine, not a new frontier model release or core research.
- Baidu
- vLLM
- Xiaomi Kunlun Suv
- DeepSeek MLA
- DeepSeek-V3.2
- Gemma4
- GLM MoE DSA
- InternVL
- Kimi-K2
- Kunlun XPU
- Llama
- Qwen3.5-MoE
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →