Users on the r/LocalLLaMA subreddit are discussing which large language models (LLMs) are performing best on M5 Max hardware. Participants are sharing their experiences with different models and quantization methods, seeking to optimize performance and token generation speed (tk/s). The conversation highlights advancements in models like Qwen and GLM, with users comparing current speeds to older benchmarks. AI
IMPACT Users are sharing insights on optimizing local LLM performance, which could inform others running similar hardware.
RANK_REASON User discussion on hardware performance with existing models.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →