PulseAugur
中
实时 10:05:32
English(EN) I pre-registered six reliability tests for my AI search measurement. None passed.

AI搜索可靠性测试在OpenAI、Gemini和Claude中均失败

对OpenAI、Gemini和Claude的AI搜索能力进行的最新分析揭示了重大的可靠性问题。从5月到9月进行的、用于测量AI搜索准确性的预先注册测试未能证实其有效性。主要发现包括每日响应不一致、测量结果之间相关性低以及许多问题缺乏变异性,这表明表面上的可靠性通常是由于问题多样性而非一致的表现。此外,观察到性能出现不明原因的下降,特别是Claude,这可能归因于搜索结果的变化或评分标准无法区分“未找到信息”与“将其与不相关数据混淆”。 AI

影响 凸显了AI搜索功能潜在的不可靠性,影响用户信任和AI驱动测量的有效性。

排序理由 对AI模型在特定测试中的性能进行分析。

在 dev.to — LLM tag 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

AI搜索可靠性测试在OpenAI、Gemini和Claude中均失败

本文如何被排名

Signal score
32 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
对AI模型在特定测试中的性能进行分析。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
product, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [2]

  1. dev.to — LLM tag TIER_1 Deutsch(DE) · Marin T. Kael ·

    我的AI搜索测量的六项预注册可靠性测试。无一通过。

    <p>Seit Mai stelle ich drei Antwortmaschinen mit Websuche (OpenAI Search, Gemini und Claude auf claude.ai) jeden Messtag dieselben 16 Fragen zu einem neuen Autor. Jede Antwort wird nach festen Regeln von minus drei bis plus drei bewertet. Am 8. Oktober erscheint der erste Band, u…

  2. dev.to — LLM tag TIER_1 English(EN) · Marin T. Kael ·

    我为我的AI搜索测量预先注册了六项可靠性测试。无一通过。

    <p>Since May I have been asking three answer engines with web search (OpenAI Search, Gemini and Claude on claude.ai) the same 16 questions about a new author. Each answer is scored by fixed rules from -3 to +3. In October the first book comes out, and the obvious next question is…