A new open reinforcement learning dataset, MiMo-V2.6-RL-oss, has been found to share a significant portion of bugs with the existing CyberGym benchmark. Both datasets draw from the ARVO collection of OSS-Fuzz bugs, with 223 identical bugs identified between them. However, a simple ID join only detects 139 of these overlaps due to renumbering of bugs in OSS-Fuzz, highlighting a potential pitfall in benchmark comparisons. The analysis indicates that while some overlap is expected, the shared bugs are slightly more than what would be expected by random chance, though it does not suggest that reported scores are inflated. AI
IMPACT Highlights potential issues in benchmark validity and reproducibility for RL agents.
RANK_REASON Analysis of overlap between two open RL datasets.
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →