PulseAugur
实时 17:50:23
English(EN) The performance came without a cost penalty: Tenet runs at $5.92 per LAB task vs $5.62 for base Kimi K3, effectively flat, while completing nearly twice as many

Fireworks AI 发布 Tenet 模型用于法律工作,在不增加成本的情况下提升性能 · 跟踪 5 个来源

Fireworks AI 推出了 Tenet,这是与 Harvey 密切合作开发的、用于长期法律工作的新模型。Tenet 在 Kimi K3 基础模型上进行了后训练,在 Legal Agent Benchmark (LAB) 等法律基准测试中表现出显著的性能提升,全通过率从 10.8% 提高到 19.7%。这一进展是在没有成本增加的情况下实现的,每个任务的成本与其基础模型相似,但完成的任务量几乎翻倍。Tenet 在未专门训练的基准测试(包括 Apex AgentsRedline Bench)上也显示出普遍的改进,并在 LegalBenchCUADMAUD 等基准测试中保持法律知识。 AI

影响 增强了法律领域的专业人工智能能力,有望提高法律文件分析和交付成果生成的效率。

排序理由 这是来自一家提供推理基础设施的公司而非前沿模型实验室的产品发布。

在 X — Fireworks (inference infra) 阅读 →

AI 生成摘要 · Google Gemini · 来自 5 个来源。 我们如何撰写摘要 →

Fireworks AI 发布 Tenet 模型用于法律工作,在不增加成本的情况下提升性能 · 跟踪 5 个来源

本文如何被排名

Signal score
16 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
这是来自一家提供推理基础设施的公司而非前沿模型实验室的产品发布。
Source corroboration
5 independent sources
Strong cross-source corroboration — multiple independent publishers covered this within the clustering window.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [5]

  1. X — Fireworks (inference infra) TIER_1 English(EN) · FireworksAI_HQ ·

    Tenet came out of a close collaboration with @ItsJulioPereyra @nikogrupen @calvincongelado @gabepereyra: many base models, recipes, and data approaches; many ru

    Tenet came out of a close collaboration with @ItsJulioPereyra @nikogrupen @calvincongelado @gabepereyra: many base models, recipes, and data approaches; many runs, rollbacks, and harness revisions to reach this checkpoint. We’re so grateful for the partnership!

  2. X — Fireworks (inference infra) TIER_1 English(EN) · FireworksAI_HQ ·

    Tenet 的性能并未带来成本的增加:每个 LAB 任务的成本为 5.92 美元,而基础 Kimi K3 为 5.62 美元,基本持平,同时完成了近两倍的任务量

    The performance came without a cost penalty: Tenet runs at $5.92 per LAB task vs $5.62 for base Kimi K3, effectively flat, while completing nearly twice as many tasks. That comes from open-weight per-token pricing and reward shaping for token efficiency.

  3. X — Fireworks (inference infra) TIER_1 English(EN) · FireworksAI_HQ ·

    收益具有普遍性。在Tenet从未训练过的基准测试中,它在@mercor的Apex Agents和@crosbylegal的Redline Bench上有所改进,后者在不同的har

    The gains generalize. On benchmarks Tenet never trained on, it improved on @mercor's Apex Agents and @crosbylegal's Redline Bench, the latter in a different harness entirely. And it showed no meaningful regression on legal knowledge benchmarks like LegalBench, CUAD and MAUD.

  4. X — Fireworks (inference infra) TIER_1 English(EN) · FireworksAI_HQ ·

    在 LAB(Harvey 的法律代理基准测试)中,代理能够生成符合数十项标准的已完成法律交付成果。

    On LAB, Harvey's Legal Agent Benchmark, agents produce finished legal deliverables graded against dozens of criteria. All-pass counts a task only if it clears them all. Tenet lifts all-pass from 10.8% to 19.7% over the Kimi K3 base, reaching state-of-the-art on LAB Contracts. h…

  5. X — Fireworks (inference infra) TIER_1 English(EN) · FireworksAI_HQ ·

    近期 @Harvey 推出了 Tenet,这是其首个为长期法律工作训练的模型。

    Recently @Harvey introduced Tenet, its first model, trained for long horizon legal work. Harvey post-trained it from a Kimi K3 base in collaboration with Fireworks using async RL on our Training API. A thread on promising initial results for both performance and https://t.co/7J…