PulseAugur
实时 11:39:15
English(EN) I analyzed 31,352 hourly LLM benchmark scores: within-day variation was 2.8 points, while between-day variation was 8.4 [P]

LLM 基准测试分数显示日间变化是日内变化的 3 倍

一项对超过 31,000 个小时的大型语言模型 (LLM) 基准测试分数的新分析揭示了模型性能的显著差异。研究发现,日内分数波动平均为 2.8 分,而日间变化则大得多,平均为 8.4 分。这表明通过观察每日趋势而非每小时波动,可以更可靠地检测到持续的性能变化,因为日间变化大约是日内变化的 3 倍。该分析是使用 AIStupidLevel 进行的,这是一个由作者开发的用于持续 LLM 基准测试和漂移检测的开源系统。 AI

影响 强调了对 LLM 性能进行稳健、持续监控的必要性,以区分真实的漂移和正常的运行方差。

排序理由 LLM 基准测试分数和性能变化分析。[lever_c_demoted from research: ic=1 ai=1.0]

在 r/MachineLearning 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

LLM 基准测试分数显示日间变化是日内变化的 3 倍

本文如何被排名

Signal score
20 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
LLM 基准测试分数和性能变化分析。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. r/MachineLearning TIER_1 English(EN) · /u/ionutvi ·

    我分析了 31,352 个小时的 LLM 基准测试分数:日内变化为 2.8 分,日间变化为 8.4 [P]

    <table> <tr><td> <a href="https://www.reddit.com/r/MachineLearning/comments/1w1jp1j/i_analyzed_31352_hourly_llm_benchmark_scores/"> <img alt="I analyzed 31,352 hourly LLM benchmark scores: within-day variation was 2.8 points, while between-day variation was 8.4 [P]" src="https://…