A user shared their experience achieving high performance with a quad setup of R9700 AI Pro GPUs using vLLM-Radiance. They reported reaching 17.6k tokens per second for prefill and a peak throughput of 106 tokens per second. This setup, running on a Gigabyte MZ32-AR0 motherboard with an EPYC 7282 CPU and the Hermes model with Qwen 3.8 27b, officially supports dual GPU configurations but the user found success with four. AI
IMPACT Demonstrates high throughput for local LLM inference, potentially improving user experience and accessibility.
RANK_REASON User-reported performance benchmark of hardware and software for local LLM deployment.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →