Sarvam AI has released its Sarvam 30B model, with performance metrics now available for several benchmarks. The model achieved 63.3% on GPQA, 7.5% on Humanity's Last Exam, and 19.2% on SciCode. Notably, it scored 0% on long-context reasoning tasks, and its cost-effectiveness is highlighted with 136.2 intelligence points per dollar. AI
IMPACT Provides benchmark data for the Sarvam 30B model, useful for comparing its capabilities against other LLMs.
RANK_REASON The item reports on benchmark results for an open-source LLM, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — fosstodon.org →
- GPQA: A Graduate-Level Google-Proof Q&A Benchmark
- Humanity's Last Exam
- long-context reasoning
- SciCode
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →