A user on Reddit's r/LocalLLaMA subreddit conducted extensive benchmarks to determine the fastest and most efficient engine for running the Qwen3.8-27B model on macOS. After five days and over 100 GPU hours of testing, the user found that MTPLX and llama.cpp with MTP (Metal Tensor Parallelism) offered the best performance for agentic coding tasks. The benchmarks involved both short synthetic tests and complex, multi-phase agentic coding challenges, as well as prefill speed tests. AI
IMPACT Identifies optimal configurations for running large language models locally, potentially improving user experience and accessibility.
RANK_REASON User-conducted benchmark comparing performance of different engines for a specific LLM on a specific OS. [lever_c_demoted from research: ic=1 ai=1.0]
- llama.cpp
- llama.cpp baseline
- llama.cpp + DFlash2
- macOS
- mlx-dspark DFlash2
- mlx-dspark DSpark
- MTPLX
- Qwen3.8-27B
- vllm-mlx
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →