A team has developed and open-sourced an optimized kernel for the Qwen3.6 35B-A3B large language model, specifically targeting AMD MI350X GPUs. Their benchmark results show that 8x MI350X GPUs can achieve over 78,000 output tokens per second, which is more than double the throughput of vLLM. This development aims to improve the performance of LLMs on AMD hardware, addressing the current dominance of NVIDIA in the market due to more mature software stacks like CUDA. AI
IMPACT Enhances LLM performance on AMD hardware, potentially offering a competitive alternative to NVIDIA GPUs for AI workloads.
RANK_REASON The cluster details an open-source software optimization for a specific LLM on particular hardware, including benchmark results. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →