A user is building a custom Local LLM Arena on a MacBook M4 to benchmark various AI models for their specific use cases. Instead of relying on internet benchmarks, they have created their own system to test five local models: Qwen3.8-27B, Ternary Bonsai 27B, GPT-OSS-20B, Gemma 4 12B, and Mistral Small 3.2 24B. The arena evaluates models on 36 tasks each, focusing on Polish language quality, reasoning, document analysis, programming, and performance metrics like speed and memory usage, all within a consistent context length of 4096. AI
IMPACT Enables personalized AI model evaluation beyond standard benchmarks.
RANK_REASON User-created tool for benchmarking AI models.
Read on Mastodon — mastodon.social →
- Codex
- Gemini Antigravity
- Gemma 4 12B
- GPT-OSS-20B
- M4 MacBook
- Mastodon
- Mistral Small 3.2 24B
- Qwen3.8-27B
- Ternary Bonsai 27B
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →