PulseAugur
中
实时 22:56:28
English(EN) [P] Pecision models that score every allowed label from the logits: Jebadiah v2.1 (27B, 9B), open weights and self-run benchmark results [P]

Jebadiah v2.1 模型发布,决策评分能力得到提升

Jebadiah v2.1,一套新的开放权重模型,已发布 27B 和 9B 参数版本。这些模型旨在将决策视为一个封闭集评分问题,而不是文本生成,从而从 logits 中为每个允许的标签评分。v2.1 更新在决策指数基准测试中显示出改进的性能,特别是在知识、语言和检索任务方面,尽管在工具使用能力方面注意到了一次回归。 AI

影响 为决策评分任务提供了新的开放权重模型,并提供了基准测试结果和代码以供进一步研究。

排序理由 发布了具有基准测试结果和代码的开放权重模型。[lever_c_demoted from research: ic=1 ai=1.0]

在 r/MachineLearning 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Jebadiah v2.1 模型发布,决策评分能力得到提升

本文如何被排名

Signal score
1 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
发布了具有基准测试结果和代码的开放权重模型。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Same-day
Cluster formed today. Ranking reflects the current source set at time of score.

完整方法见我们的编辑标准。

报道来源 [1]

  1. r/MachineLearning TIER_1 English(EN) · /u/WebDevToday ·

    Jebadiah v2.1 (27B, 9B) 的精确模型可以从 logits 中为每个允许的标签评分:开放权重和自运行基准测试结果

    <!-- SC_OFF --><div class="md"><p>I've been building open models that treat a decision as a closed-set scoring problem rather than text generation. The input is structured context plus a typed question with a fixed set of options. The output is a probability for each option, take…