PulseAugur
中
实时 17:18:41
English(EN) An open RL dataset (MiMo-V2.6-RL-oss, cyber) shares 223 of CyberGym's 1,507 bugs; both draw on ARVO. A plain ID join finds only 139, because OSS-Fuzz renumbered

开放的强化学习数据集共享bug,使基准测试比较复杂化

新发布的开放强化学习数据集MiMo-V2.6-RL-oss被发现与现有的CyberGym基准测试共享了相当一部分bug。这两个数据集都借鉴了OSS-Fuzz bug的ARVO集合,两者之间识别出223个相同的bug。然而,由于OSS-Fuzz中的bug重新编号,简单的ID连接只能检测到139个重叠,这凸显了基准测试比较中潜在的陷阱。分析表明,虽然一些重叠是预料之中的,但共享的bug数量略多于随机概率的预期,尽管这并不表明报告的分数被夸大了。 AI

影响 突出了强化学习代理基准测试有效性和可复现性方面潜在的问题。

排序理由 对两个开放强化学习数据集之间重叠的分析。

在 Mastodon — mastodon.social 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

开放的强化学习数据集共享bug,使基准测试比较复杂化

本文如何被排名

Signal score
15 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
对两个开放强化学习数据集之间重叠的分析。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [2]

  1. dev.to — LLM tag TIER_1 English(EN) · Raimondas Lencevicius ·

    相同错误,两个ID:一个开放RL数据集共享CyberGym 15%的错误,而简单的ID连接错过了38%

    <p>Open RL environments are great for research. They also open another route for benchmark overlap. Here is a small, reproducible case, plus one pitfall that will bite anyone checking for it.</p> <p>(Short version: this does not show that any reported score is inflated; caveats b…

  2. Mastodon — mastodon.social TIER_1 English(EN) · [email protected] ·

    一个开放的RL数据集(MiMo-V2.6-RL-oss, cyber)分享了CyberGym的1507个bug中的223个;两者都借鉴了ARVO。简单的ID连接只找到了139个,因为OSS-Fuzz重新编号了

    An open RL dataset (MiMo-V2.6-RL-oss, cyber) shares 223 of CyberGym's 1,507 bugs; both draw on ARVO. A plain ID join finds only 139, because OSS-Fuzz renumbered its bugs, and a 13-gram filter against CyberGym's task descriptions flags none. Task rows hold no PoCs; this doesn't sh…