The GLM-5.2 model, a 744B-parameter Mixture of Experts (MoE) model, is reportedly runnable on consumer hardware. One user shared that it can operate on a machine with approximately 25 GB of RAM, while another detailed its performance on a MacBook Pro M5 with 48 GB of RAM, achieving speeds between 2 to 2.8 tokens per second. This development suggests increased accessibility for running large language models locally. AI
IMPACT Enables wider accessibility and experimentation with large language models on personal devices.
RANK_REASON User-reported performance and accessibility of a large language model on consumer hardware.
AI-generated summary · Google Gemini · from 3 sources. How we write summaries →