PulseAugur
实时 01:29:12
English(EN) Optimal control grows neural network layers where error peaks A new arXiv preprint reframes neural network training as optimal control, inserting layers where e

OpenAI 标记 SWE-bench Pro 编码基准测试缺陷;新研究重构神经网络训练

OpenAI 发现其先前曾认可为被污染的 SWE-bench 基准测试的替代品的 SWE-bench Pro 编码基准测试存在严重的可靠性问题。这项新分析表明,SWE-bench Pro 本身可能不是衡量 AI 编码性能的可靠指标。另外,一篇新的 arXiv 预印本提出将神经网络训练重构为最优控制问题,旨在误差最高处插入层数,但目前支持该方法的证据有限。 AI

影响 对 AI 评估基准测试的批评和新颖的训练方法突显了在可靠评估和改进 AI 能力方面持续存在的挑战。

排序理由 该集群包含一篇研究论文和一个对基准测试的批评,符合研究类别。

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

OpenAI 标记 SWE-bench Pro 编码基准测试缺陷;新研究重构神经网络训练

本文如何被排名

Signal score
14 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群包含一篇研究论文和一个对基准测试的批评,符合研究类别。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [2]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    OpenAI发现SWE-bench Pro编码基准测试存在缺陷 OpenAI的新分析揭示了SWE-bench Pro的可靠性问题——正是它曾推荐用于重新评估的模型

    OpenAI finds flaws in SWE-bench Pro coding benchmark OpenAI's new analysis exposes reliability issues in SWE-bench Pro — the very benchmark it recommended to replace the contaminated SWE-bench Verified — leaving A https://www. notatechguy.com/openai-finds-f laws-in-swe-bench-pro-…

  2. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    最优控制在误差达到峰值处增长神经网络层数 一篇新的arXiv预印本将神经网络训练重构为最优控制,在误差达到峰值处插入层数

    Optimal control grows neural network layers where error peaks A new arXiv preprint reframes neural network training as optimal control, inserting layers where error is highest — but evidence is limited to scientific datase https://www. notatechguy.com/optimal-contro l-grows-neura…