PulseAugur
实时 01:06:04
English(EN) Grok 4.6's most important number is one xAI didn't even advertise

Grok 4.6 显示出重大校准改进,xAI 未宣传 · 跟踪到 2 个来源

xAIGrok 4.6 模型在不幻觉率方面有了显著提高,从 45.9% 提高到 65.7%。这一指标衡量的是模型在不确定时倾向于不回答而不是捏造回应,xAI 在官方公告中并未强调这一点。虽然在智能和编码方面的头条基准显示 Grok 4.6 与 GPT-5.6 Sol 等竞争对手相当,但其增强的校准对于需要多步错误累积的代理任务至关重要。这一改进表明,尽管 Sol 在一些高级代理基准测试中仍处于领先地位,但 Grok 4.6 在复杂的决策过程中可能更可靠。 AI

影响 通过减少捏造的回应,提高了代理任务的可靠性,可能使其更适合复杂的多步决策。

排序理由 集群讨论了来自前沿实验室 (xAI) 的新模型发布 (Grok 4.6),并附有具体的基准数据。[lever_c_demoted from frontier_release: ic=2 ai=1.0]

在 r/cursor 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

Grok 4.6 显示出重大校准改进,xAI 未宣传 · 跟踪到 2 个来源

报道来源 [2]

  1. r/cursor TIER_2 English(EN) · /u/Kai_ThoughtArchitect ·

    Grok 4.6 最重要的数字是 xAI 甚至没有宣传的一个

    <!-- SC_OFF --><div class="md"><p>Grok 4.6 dropped yesterday and the debate is the usual &quot;is it better than Sol?&quot; On the headline benchmarks it's a genuine tie: Intelligence Index 61 vs 61, Coding 76.8 vs 77.4, Agentic 58.7 vs 57.8.</p> <p>The number nobody screenshots …

  2. r/singularity TIER_2 English(EN) · /u/Kai_ThoughtArchitect ·

    Grok 4.6 最重要的数字是 xAI 甚至没有宣传的一个

    <!-- SC_OFF --><div class="md"><p>Grok 4.6 dropped yesterday and the debate is the usual &quot;is it better than Sol?&quot; On the headline benchmarks it's a genuine tie: Intelligence Index 61 vs 61, Coding 76.8 vs 77.4, Agentic 58.7 vs 57.8.</p> <p>The number nobody screenshots …