PulseAugur
EN
LIVE 00:04:07

Ling 3.0 language model shows impressive speed on Strix Halo hardware

Ling 3.0, a new iteration of the Ling language model, has been demonstrated running on the Strix Halo hardware. This version utilizes vLLM with ROCm/HiP and 4-bit compressed tensors (int4) for enhanced performance. Early comparisons suggest that Ling 3.0 significantly outperforms Qwen-122b in speed when both are optimized for their respective formats, though tool call functionality is noted as being broken in certain harnesses. AI

IMPACT Demonstrates potential for faster inference on specialized hardware, impacting local LLM deployment.

RANK_REASON This is a research-level update on a specific language model's performance on new hardware, not a frontier release from a major lab. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Ling 3.0 language model shows impressive speed on Strix Halo hardware

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Badger-Purple ·

    Ling 3.0 Flash on Strix Halo

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1vkz5do/ling_30_flash_on_strix_halo/"> <img alt="Ling 3.0 Flash on Strix Halo" src="https://preview.redd.it/ay0qneprgmih1.jpeg?width=640&amp;crop=smart&amp;auto=webp&amp;s=04f025fe46e0277327ee686dca8a9e759d913…