A user has created a custom benchmark system called "Local LLM Arena" to evaluate the performance of five different large language models (LLMs) running locally on their M4 MacBook with 16GB of unified memory. The goal is to determine which models are most practical for everyday use on this specific hardware, rather than relying on theoretical benchmarks. The tested models include Gemma 4 12B, Ternary Bonsai 27B, GPT-OSS-20B, Qwen3.8-27B, and Mistral Small 3.2 24B, with each model undergoing 36 tasks for a total of 180 tests. The benchmark assesses various aspects such as Polish language quality, reasoning, document analysis, programming, and agent tasks, while also measuring speed, memory usage, and stability. AI
IMPACT Provides practical insights into local LLM performance on consumer hardware, guiding user choices for on-device AI applications.
RANK_REASON User-created benchmark system for evaluating LLMs on personal hardware.
Read on Mastodon — sigmoid.social →
- Codex
- Gemini Antigravity
- Gemma 4 12B
- GPT-OSS-20B
- M4 MacBook
- Mastodon
- Mistral Small 3.2 24B
- Qwen3.8-27B
- Ternary Bonsai 27B
- MacBook
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →