PulseAugur
EN
LIVE 22:55:21

Swift-1.5-Qwen3.8-27b-oQ8e-mtp benchmarks on Apple M5 Max

A user benchmarked the Swift-1.5-Qwen3.8-27b-oQ8e-mtp model on an Apple M5 Max chip, achieving a generation speed of 34.6 tokens/second. The Swift model produced significantly fewer tokens on average across various scenarios like code generation and agent workflows compared to the base Qwen 3.8 27B model, resulting in shorter benchmark run times. Despite generating less output, the quality score was comparable, suggesting Swift-1.5-Qwen3.8-27b-oQ8e-mtp could be a viable daily driver for its efficiency. AI

IMPACT Provides performance data for local LLM deployment on Apple hardware, aiding users in selecting efficient models.

RANK_REASON User benchmark of a specific model variant on consumer hardware.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Swift-1.5-Qwen3.8-27b-oQ8e-mtp benchmarks on Apple M5 Max

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/DerTomsn ·

    Swift-1.5-Qwen3.8-27b-oQ8e-mtp on Apple M5 Max — 34.8 tok/s — llm-bench.io

    <table> <tr><td> <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wte7n0/swift15qwen3827boq8emtp_on_apple_m5_max_348_toks/"> <img alt="Swift-1.5-Qwen3.8-27b-oQ8e-mtp on Apple M5 Max — 34.8 tok/s — llm-bench.io" src="https://external-preview.redd.it/Ph4uY4Gy5DogykJmTAQxXAzE2…