A user on Reddit shared their experience running the Qwen 3.8 27B model on a MacBook Max with 128GB of RAM. They reported initial speeds of around 30 tokens/second, which only slightly decreased to 27 tokens/second even with a context window usage of up to 50,000 tokens. The user is seeking advice on optimizing performance and confirming their quantization settings for the model, which is running via Unsloth Studio. AI
IMPACT Provides insights into the practical performance of large language models on high-end consumer hardware.
RANK_REASON User-generated performance test of a specific model on consumer hardware.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →