PulseAugur
中
实时 15:02:56
English(EN) StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

StateM 运行时通过 Harness Scaling 提高 AI 代理准确性 · 跟踪 2 个来源

研究人员开发了 StateM,一个旨在增强长时程 AI 代理性能而无需修改其底层模型权重的新运行时系统。该系统围绕持久化状态、可恢复的运行手册和可强制执行的过程控制来组织代理执行。StateM 在 Terminal-Bench 2.1 等基准测试中展示了显著的改进,将 GPT-5.5 xhigh 的准确率提高到 92.1%,并使用 GPT-5.6 Sol xhigh 达到了 95.3% 的原始准确率。它还以极低的适应成本将 DeepSeek-V4 Flash 的准确率从 82.7% 提高到 88.1%。 AI

影响 增强了长时程代理的能力并降低了推理成本,可能加速其在复杂任务自动化中的应用。

排序理由 该集群描述了一篇详细介绍用于提高 AI 代理性能的新颖系统的研究论文。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

StateM 运行时通过 Harness Scaling 提高 AI 代理准确性 · 跟踪 2 个来源

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群描述了一篇详细介绍用于提高 AI 代理性能的新颖系统的研究论文。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
53 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [2]

  1. arXiv cs.AI TIER_1 English(EN) · Ziheng Qin, Yaxin Lu, Zhangyang Atlas Wang, Kai Wang ·

    StateM:通过Harness Scaling在Terminal-Bench 2.1上达到95.3%的原始准确率,或15美元的Frontier运行

    arXiv:2608.15089v1 Announce Type: new Abstract: Long-horizon agents can fail even when their underlying models can solve the constituent steps. They may lose track of mutable state, fail to reactivate lessons from earlier executions, skip known procedures, or stop prematurely. We…

  2. Hugging Face Daily Papers TIER_1 English(EN) ·

    StateM:通过Harness Scaling在Terminal-Bench 2.1上达到95.3%的原始准确率,或15美元的Frontier运行

    StateM is a runtime system that improves long-horizon agent execution through durable states, recoverable runbooks, and enforceable procedural controls without altering model weights.