PulseAugur
EN
LIVE 19:31:48

GLM-5.2 runs at 12.2 tok/s on 16x AMD MI50 GPUs via llama.cpp

A user on r/LocalLLaMA shared their experience running the GLM-5.2 model with a Q4 quantization using llama.cpp RPC. They achieved an output speed of 12.2 tokens per second and an input speed of 30.9 tokens per second on a document with 10.7k tokens. This setup utilized 16 AMD MI50 GPUs with 32GB of VRAM each, totaling 512GB across two nodes, and a 10 GbE interconnect. AI

IMPACT Demonstrates achievable inference speeds for large language models on consumer/prosumer hardware, informing infrastructure choices.

RANK_REASON User-reported performance of a specific model on specific hardware using a specific software framework.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

GLM-5.2 runs at 12.2 tok/s on 16x AMD MI50 GPUs via llama.cpp

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Legal-Ad-3901 ·

    16x AMD MI50 32GB: GLM-5.2 Q4 at 12.2 tok/s with llama.cpp RPC

    <!-- SC_OFF --><div class="md"><p>GLM-5.2 UD-Q4_K_XL GGUF @ 12.2 tok/s output // 30.9 tok/s input on a real 10.7k-token document using llama.cpp RPC - At 10.7k context: 10.2 tok/s output with coherent long-form generation</p> <ul> <li>Two parallel requests: 14.5 tok/s aggregate</…