PulseAugur
实时 07:05:47
English(EN) Benchmark Contamination: A Taxonomy Organized by Defeated Mitigation

新的分类法按被挫败的缓解措施对人工智能基准污染进行分类

提出了一种新的人工智能基准污染分类法,该分类法按其挫败的缓解方法对污染类型进行组织。该分类法将污染分为直接型、派生型、时间型、分布型和习得型,并解决了训练时和评估时的数据泄露问题。研究还引入了一个四字段披露协议来报告污染状态,并承认“未知”是一个有效条目。对41份文件的分析显示,一些变量的编码者间可靠性较低,在变量适用时间上存在显著分歧,而不是在其陈述内容上存在分歧。 AI

影响 这项研究旨在提高人工智能基准报告的可靠性和透明度,从而可能更准确地评估模型能力。

排序理由 该集群包含一篇学术论文,详细介绍了人工智能基准污染的新分类法。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的分类法按被挫败的缓解措施对人工智能基准污染进行分类

本文如何被排名

Signal score
25 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群包含一篇学术论文,详细介绍了人工智能基准污染的新分类法。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Johanna Angulo, V\'ictor Yeste, Hector Espinos-Morato ·

    基准测试污染:按被击败的缓解措施分类的分类法

    arXiv:2608.29463v1 Announce Type: cross Abstract: A benchmark score is a joint property of the model, the evaluation harness, the elicitation budget, the sampled population, and contamination status. Leaderboards publish the model and the score, so capability and leakage stay obs…