PulseAugur
EN
LIVE 18:29:05

MTPLX and llama.cpp+MTP lead macOS benchmarks for Qwen3.8-27B

A user on Reddit's r/LocalLLaMA subreddit conducted extensive benchmarks to determine the fastest and most efficient engine for running the Qwen3.8-27B model on macOS. After five days and over 100 GPU hours of testing, the user found that MTPLX and llama.cpp with MTP (Metal Tensor Parallelism) offered the best performance for agentic coding tasks. The benchmarks involved both short synthetic tests and complex, multi-phase agentic coding challenges, as well as prefill speed tests. AI

IMPACT Identifies optimal configurations for running large language models locally, potentially improving user experience and accessibility.

RANK_REASON User-conducted benchmark comparing performance of different engines for a specific LLM on a specific OS. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

MTPLX and llama.cpp+MTP lead macOS benchmarks for Qwen3.8-27B

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/ex-arman68 ·

    Benchmark results: what is the best and fastest engine to run Qwen3.8-27B on macOS

    <!-- SC_OFF --><div class="md"><p>The new Qwen 3.8 27B is fantastic for local agentic use. The problem is, what makes it so good, being a dense model, also makes it slow. Many engines and versions of the model claim various speed increase. How true are those claim? And does a pro…