A user on Reddit's r/LocalLLaMA subreddit reported a significant performance increase when switching from the Qwen3.8 27B model to Ornith 1.5 35B-A3B for local agent tasks. The Ornith model achieved approximately 180 tokens/second, a threefold improvement over the Qwen model's 60 tokens/second, while maintaining comparable performance on the user's custom tests for tool use and coding. This speed boost is attributed to Ornith's architecture, which uses a smaller number of active parameters per token and a hybrid attention mechanism that limits KV cache growth, allowing for larger context windows with minimal performance degradation. AI
IMPACT Offers a significant speed improvement for local AI agent operations, potentially enabling more complex tasks on consumer hardware.
RANK_REASON User-reported performance comparison of local LLMs for agent tasks.
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →