A user on Reddit shared benchmarks for the EschaLabs/Qwen3.6-35B-A3B-Escha-W2 model, comparing it against the APEX (Q5 Balanced) model. The Escha model demonstrated significantly faster generation and prefill speeds, being 1.85x and 2.48x faster respectively. While the APEX model showed a lower perplexity on wikitext-2, the Escha model performed equally or better on instruction adherence, math reasoning, code generation, and PhD-level questions, with a notable win on GPQA-Diamond. AI
IMPACT Demonstrates potential for efficient, high-performing local LLMs, especially for users with AMD GPUs.
RANK_REASON User-generated benchmark results for a specific model variant. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →