PulseAugur
中
实时 10:05:47
English(EN) Switched my local agent from Qwen3.8 27B to Ornith 1.5 35B-A3B on two 5070 Tis: about 180 tok/s vs 60, same scores on my tests

Ornith 1.5 35B-A3B 模型为本地 AI 代理提供 3 倍速度提升

Reddit r/LocalLLaMA 版块的一名用户报告称,在将 Qwen3.8 27B 模型切换到 Ornith 1.5 35B-A3B 进行本地代理任务时,性能显著提升。Ornith 模型达到了大约 180 tokens/秒,比 Qwen 模型 60 tokens/秒的速度提升了三倍,同时在用户用于工具使用和编码的自定义测试中保持了可比的性能。这种速度提升归因于 Ornith 的架构,该架构每个 token 使用的激活参数数量较少,并采用混合注意力机制限制 KV 缓存增长,从而在性能几乎没有下降的情况下实现更大的上下文窗口。 AI

影响 为本地 AI 代理操作提供了显著的速度提升,有可能在消费级硬件上实现更复杂的任务。

排序理由 用户报告的用于代理任务的本地 LLM 性能对比。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Ornith 1.5 35B-A3B 模型为本地 AI 代理提供 3 倍速度提升

本文如何被排名

Signal score
8 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
用户报告的用于代理任务的本地 LLM 性能对比。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Excellent-Issue-5956 ·

    将本地代理从 Qwen3.8 27B 切换到 Ornith 1.5 35B-A3B(两块 5070 Ti):约 180 token/秒 对比 60 token/秒,我的测试得分相同

    <!-- SC_OFF --><div class="md"><p>My setup is two RTX 5070 Ti 16GB cards (the second one is on an OCuLink dock) with 64GB of RAM, Ollama on Windows, and the agent runs on pi in WSL. Until last night the daily model was Qwen3.8 27B UD-Q4_K_XL at 128K with MTP, which does about 55 …