PulseAugur
中
实时 13:15:47
한국어(KO) 허깅페이스 리더보드 1위 9개: 오픈소스 LLM Darwin-180B-RSI와 재귀적 자가개선(RSI) 개발자 가이드

Darwin-180B-RSI凭借递归自改进在9项Hugging Face基准测试中领先

VIDRAFT的Darwin-180B-RSI模型家族在9项官方Hugging Face基准测试中均取得第一名,超越了所有其他参与组织。该开源模型采用了递归自改进技术,即它解决可验证的问题,只保留正确的解决方案,并在没有人类生成答案的情况下进行再训练。值得注意的是,它在结构化输出生成(IFStruct)和从PDF提取数据(ExtractBench)等实际任务中表现出色,展示了其在现实世界中的应用潜力。 AI

影响 为开源模型在结构化数据任务和基准测试上的性能树立了新标杆。

排序理由 该集群详细介绍了一款新的开源模型发布及其在基准测试上的表现,包括一种新颖的训练技术。

在 dev.to — LLM tag 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

Darwin-180B-RSI凭借递归自改进在9项Hugging Face基准测试中领先

本文如何被排名

Signal score
23 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Research
该集群详细介绍了一款新的开源模型发布及其在基准测试上的表现,包括一种新颖的训练技术。
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准。

报道来源 [2]

  1. dev.to — LLM tag TIER_1 Español(ES) · 김민식/학생 ·

    Darwin-180B-RSI:Hugging Face 9项官方榜单中的头号开源语言模型

    <h2> TL;DR </h2> <ul> <li>La familia <strong>Darwin-180B-RSI</strong> de VIDRAFT ocupa el <strong>primer lugar en 9 de los 48 benchmarks oficiales</strong> de Hugging Face. Ninguna de las 95 organizaciones participantes tiene más; detrás vienen Moonshot AI y Zhipu AI (Z.ai) con 4…

  2. dev.to — LLM tag TIER_1 한국어(KO) · Vidraft_lab ·

    Hugging Face 排行榜前九名:开源 LLM Darwin-180B-RSI 与递归自我改进 (RSI) 开发者指南

    <h2> TL;DR </h2> <ul> <li>비드래프트(VIDRAFT)의 오픈소스 LLM <strong>Darwin-180B-RSI</strong> 계열이 허깅페이스 공인(official) 벤치마크 48개 중 <strong>9개에서 1위</strong>입니다. 95개 조직 가운데 가장 많고, 다음은 문샷AI·지푸AI(각 4개)입니다.</li> <li>새로 추가된 1위는 실무형 두 개입니다. <strong>IFStruct 98.95%</strong>(구조화 출력, 2,000문항 중 1,979 통과…