PulseAugur
EN
LIVE 03:30:13
Español(ES) De los exámenes al trabajo real: un modelo abierto de 180B lidera 10 clasificaciones oficiales de Hugging Face

Open 180B Model Darwin-180B-RSI Leads 10 Hugging Face Benchmarks

An open-weight model named Darwin-180B-RSI has achieved the top position on 10 official Hugging Face leaderboards, surpassing all other participating organizations. This model, built by VIDRAFT on Alibaba's Qwen3.8-Flash-Next, demonstrates strong performance not only in academic benchmarks like MMLU-Pro and GPQA but also in practical, enterprise-focused tasks such as document parsing (MDPBench), structured data output (IFStruct), and field extraction from PDFs (ExtractBench). The model's training methodology, known as recursive self-improvement (RSI), allows it to refine its capabilities by solving problems, verifying its own solutions, and retraining on confirmed correct outputs without human-labeled data. AI

IMPACT Sets a new standard for open-weight models in both academic and practical enterprise tasks, potentially accelerating adoption of advanced AI agents.

RANK_REASON Open-weight model release with multiple benchmark wins and details on training methodology. [lever_c_demoted from frontier_release: ic=2 ai=1.0]

Read on dev.to — LLM tag →

AI-generated summary · Google Gemini · from 2 sources. How we write summaries →

Open 180B Model Darwin-180B-RSI Leads 10 Hugging Face Benchmarks

How we ranked this

Signal score
54 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Significant
Open-weight model release with multiple benchmark wins and details on training methodology. [lever_c_demoted from frontier_release: ic=2 ai=1.0]
Source corroboration
2 independent sources
Multiple independent publishers reporting the same story raises confidence that it's real and newsworthy.
Topics
model release, product
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [2]

  1. dev.to — LLM tag TIER_1 Español(ES) · 김민식/학생 ·

    From exams to real work: an open 180B model leads 10 official Hugging Face rankings

    <h2> TL;DR </h2> <ul> <li>Un modelo de código abierto, <strong>Darwin-180B-RSI</strong>, ocupa ahora el <strong>puesto #1 en 10 de los 48 benchmarks oficiales de Hugging Face</strong>, el mayor número entre las 95 organizaciones que compiten. La segunda posición es Z.ai, con 4.</…

  2. dev.to — LLM tag TIER_1 English(EN) · AI OpenFree ·

    Open 180B Model Leads 10 Official Hugging Face Leaderboards, and the Zero-Token Judge Behind It

    <p><a class="article-body-image-wrapper" href="https://media2.dev.to/dynamic/image/width=800%2Cheight=%2Cfit=scale-down%2Cgravity=auto%2Cformat=auto/https%3A%2F%2Fquickchart.io%2Fchart%3Fw%3D1000%26h%3D420%26c%3D%257B%2522type%2522%253A%2520%2522horizontalBar%2522%252C%2520%2522d…