PulseAugur
EN
LIVE 16:20:58

Apple M4 Max Mac Studio leads local AI inference on decode throughput

Apple's M4 Max chip, featured in the Mac Studio, demonstrates strong local AI performance, particularly in decode throughput, outperforming competitors like NVIDIA's GB10 and AMD's Strix Halo. This advantage stems from its high memory bandwidth, crucial for sequential LLM inference where streaming model weights from memory to the GPU is a bottleneck. While Apple Silicon offers a compelling combination of large unified memory and high bandwidth, availability and configuration options for the Mac Studio have become more limited. AI

IMPACT Sets a new benchmark for local AI inference performance, particularly in decode throughput, potentially influencing future hardware design and user expectations for on-device AI capabilities.

RANK_REASON Comparison of AI hardware performance on specific benchmarks. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Tom's Hardware →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Apple M4 Max Mac Studio leads local AI inference on decode throughput

COVERAGE [1]

  1. Tom's Hardware TIER_1 English(EN) · Jeffrey Kampman ·

    Exploring Apple Silicon’s local AI performance with the Mac Studio and M4 Max — M4 Max beats GB10 and Strix Halo in decode throughput, but memory bandwidth isn't everything

    Apple Silicon has been a popular choice for local AI exploration thanks to its high memory bandwidth compared to other unified memory platforms. We tested the M4 Max version of Apple's Mac Studio to see whether its 546GB/s of bandwidth makes it the clear winner in local LLM infer…