PulseAugur
中
实时 04:00:57
English(EN) Local LLM 35B MoE — Real-world coding benchmarks (Qwen vs Ornith vs KAT)

KAT Coder 2.5 Dev 在真实编码基准测试中表现优于 Qwen 和 Ornith

一位用户对几款 35B MoE 类的本地大型语言模型进行了真实编码基准测试,比较了 Qwen 3.6、Ornith 1.0 和 KAT Coder 2.5 Dev。虽然 Qwen 3.6 作为基线模型表现良好,但在较长任务中表现出自信幻觉和漂移的倾向。Ornith 1.0 尽管基准测试表现出色,但经常过度思考简单任务,并表现出与 Qwen 类似的推理循环问题。然而,KAT Coder 2.5 Dev 以更果断的输出、在实际编码场景中更好的表现以及更少的重复性问题给用户带来了惊喜,这标志着本地模型在严肃编码工作负载方面的实用性向前迈进了一步。 AI

影响 强调 KAT Coder 2.5 Dev 是一个有前途的本地模型,适用于严肃的编码任务,可能影响开发人员工具的选择。

排序理由 用户生成的用于编码任务的本地 LLM 比较,而非主要发布或研究论文。

在 r/LocalLLaMA 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

KAT Coder 2.5 Dev 在真实编码基准测试中表现优于 Qwen 和 Ornith

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Commentary
用户生成的用于编码任务的本地 LLM 比较,而非主要发布或研究论文。
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
product, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
56 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/Undici77 ·

    本地大模型 35B MoE — 真实代码基准测试 (Qwen vs Ornith vs KAT)

    <!-- SC_OFF --><div class="md"><p>I’ve been running a fairly opinionated evaluation loop on ~35B A3B/MoE-class models for coding over the past few months. Not synthetic benchmarks: actual dev workflows, iterative debugging, refactoring passes, and failure recovery.</p> <p>Here’s …