PulseAugur
实时 13:50:50
한국어(KO) Artificial Analysis가 Harvey의 Legal Agent Benchmark(LAB)를 Stirrup 에이전트로 구현한 Harvey LAB‑AA 결과를 공개했습니다. Harvey 데이터셋(120개 비공개 과제, 24개 법률 분야)으로 기준별 합격률과 전항목(all-pass

Kimi K3 以 94.6% 的通过率领跑 Harvey 法律代理基准

Artificial Analysis 发布了其 Harvey LAB-AA 基准测试的结果,该测试使用 Stirrup 代理实现了 Harvey 法律代理基准 (LAB)。评估使用了 Harvey 数据集,该数据集包含 24 个法律领域的 120 个私有任务,以评估按标准划分的通过率和总体通过率。Kimi K3 取得了最高分 94.6%,其次是 Claude Fable 5 的 93.6% 和 Muse Spark 1.1 的 93.1%。结果还包括成本、代币和速度的指标。 AI

影响 为法律任务中的 AI 代理设定了新的基准,突出了 Kimi K3 的领先表现,并为其他模型提供了比较数据。

排序理由 该集群报告了 AI 代理在法律任务上的基准评估结果。[lever_c_demoted from research: ic=1 ai=1.0]

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Kimi K3 以 94.6% 的通过率领跑 Harvey 法律代理基准

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群报告了 AI 代理在法律任务上的基准评估结果。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
product, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
49 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准

报道来源 [1]

  1. Mastodon — fosstodon.org TIER_1 한국어(KO) · [email protected] ·

    Artificial Analysis 发布了 Harvey LAB-AA 的结果,这是 Harvey 法律代理基准 (LAB) 在 Stirrup 代理上的实现。Harvey 数据集(120 个私有任务,24 个法律领域)显示了每个基准的通过率和总通过率。

    Artificial Analysis가 Harvey의 Legal Agent Benchmark(LAB)를 Stirrup 에이전트로 구현한 Harvey LAB‑AA 결과를 공개했습니다. Harvey 데이터셋(120개 비공개 과제, 24개 법률 분야)으로 기준별 합격률과 전항목(all-pass) 합격률을 평가했으며, Kimi K3가 94.6%로 1위, Claude Fable 5 93.6%, Muse Spark 1.1 93.1%로 뒤를 이었습니다. 비용·토큰·속도 지표도 함께 제시됩니다. https:// …