A user on r/LocalLLaMA shared their experience running the GLM-5.2 model with a Q4 quantization using llama.cpp RPC. They achieved an output speed of 12.2 tokens per second and an input speed of 30.9 tokens per second on a document with 10.7k tokens. This setup utilized 16 AMD MI50 GPUs with 32GB of VRAM each, totaling 512GB across two nodes, and a 10 GbE interconnect. AI
IMPACT Demonstrates achievable inference speeds for large language models on consumer/prosumer hardware, informing infrastructure choices.
RANK_REASON User-reported performance of a specific model on specific hardware using a specific software framework.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →