Ling 3.0, a new iteration of the Ling language model, has been demonstrated running on the Strix Halo hardware. This version utilizes vLLM with ROCm/HiP and 4-bit compressed tensors (int4) for enhanced performance. Early comparisons suggest that Ling 3.0 significantly outperforms Qwen-122b in speed when both are optimized for their respective formats, though tool call functionality is noted as being broken in certain harnesses. AI
IMPACT Demonstrates potential for faster inference on specialized hardware, impacting local LLM deployment.
RANK_REASON This is a research-level update on a specific language model's performance on new hardware, not a frontier release from a major lab. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →