PulseAugur
EN
LIVE 09:44:17

Ornith 1.5 35B-A3B model offers 3x speed boost for local AI agents

A user on Reddit's r/LocalLLaMA subreddit reported a significant performance increase when switching from the Qwen3.8 27B model to Ornith 1.5 35B-A3B for local agent tasks. The Ornith model achieved approximately 180 tokens/second, a threefold improvement over the Qwen model's 60 tokens/second, while maintaining comparable performance on the user's custom tests for tool use and coding. This speed boost is attributed to Ornith's architecture, which uses a smaller number of active parameters per token and a hybrid attention mechanism that limits KV cache growth, allowing for larger context windows with minimal performance degradation. AI

IMPACT Offers a significant speed improvement for local AI agent operations, potentially enabling more complex tasks on consumer hardware.

RANK_REASON User-reported performance comparison of local LLMs for agent tasks.

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Ornith 1.5 35B-A3B model offers 3x speed boost for local AI agents

How we ranked this

Signal score
9 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
User-reported performance comparison of local LLMs for agent tasks.
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Excellent-Issue-5956 ·

    Switched my local agent from Qwen3.8 27B to Ornith 1.5 35B-A3B on two 5070 Tis: about 180 tok/s vs 60, same scores on my tests

    <!-- SC_OFF --><div class="md"><p>My setup is two RTX 5070 Ti 16GB cards (the second one is on an OCuLink dock) with 64GB of RAM, Ollama on Windows, and the agent runs on pi in WSL. Until last night the daily model was Qwen3.8 27B UD-Q4_K_XL at 128K with MTP, which does about 55 …