PulseAugur
实时 13:52:21
English(EN) So what benchmarks are AI companies using internally?

在外部测试担忧之际,AI公司的内部基准测试受到质疑

讨论围绕着AI基准测试的有效性和潜在操纵性展开。用户质疑像OpenAI和Anthropic这样的大型AI公司是否使用内部、不太公开的基准测试来追踪模型的真正进展,因为外部基准测试被认为很容易被操纵。对话强调了对超越公开可用测试来评估AI能力的更可靠方法的需求。 AI

影响 引发了对AI性能指标和内部评估方法可靠性的质疑。

排序理由 用户生成的关于AI开发实践相关主题的讨论。

在 r/OpenAI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

在外部测试担忧之际,AI公司的内部基准测试受到质疑

报道来源 [1]

  1. r/OpenAI TIER_2 English(EN) · /u/Rusofil__ ·

    So what benchmarks are AI companies using internally?

    <!-- SC_OFF --><div class="md"><p>We're all familiar with benchmaxing and how it's not valuable measurement on AI's capability. So there must be some internal tests openai, anthropic and others are using internally to track real progress of their models that are not skewed by try…