The third iteration of the Local LLM Arena benchmark tested five models on a 16GB MacBook M4. GPT-OSS-20B emerged as the top performer overall, offering strong reasoning capabilities and good performance in Polish and German translation, achieving approximately 20 tokens/sec. Qwen 27B excelled in Polish language tasks and document understanding but was significantly slower at around 3 tokens/sec. Mistral 24B demonstrated stability with a perfect score across all tests, while Gemma 12B was fast but frequently produced empty responses. Bonsai 27B also achieved a perfect score but was deemed the weakest overall. AI
IMPACT Provides comparative performance data for running local LLMs on consumer hardware, aiding developers in model selection.
RANK_REASON Benchmark results for multiple local LLMs on consumer hardware. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →