A Reddit user shared benchmarks for a dual Radeon MI50 setup, utilizing two GPUs each with 16GB of HBM2 VRAM. The user tested various dense and Mixture of Experts (MoE) models from Huggingface, aiming to stay within the 32GB total VRAM limit. The benchmarks were run using llama.cpp on Ubuntu with Vulkan, and the results show varying performance across different models and quantization levels, with some MoE models demonstrating significantly higher throughput. AI
IMPACT Provides performance data for running LLMs on specific AMD GPU hardware, informing users about potential throughput.
RANK_REASON User-generated benchmarks for specific hardware and software configurations.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →