PulseAugur
实时 13:16:51
English(EN) 📊 Sarvam 30B (high) — the actual numbers GPQA: 63.3% Humanity's Last Exam: 7.5% Long Context Reasoning: 0% SciCode: 19.2% 💰 136.2 intelligence points per dollar

Sarvam 30B 模型性能指标在基准测试中公布

Sarvam AI 发布了其 Sarvam 30B 模型,目前已公布多个基准测试的性能指标。该模型在 GPQA 上达到 63.3%,在 Humanity's Last Exam 上达到 7.5%,在 SciCode 上达到 19.2%。值得注意的是,它在长上下文推理任务上得分为 0%,并且其成本效益以每美元 136.2 智能点突出显示。 AI

影响 提供了 Sarvam 30B 模型的基准数据,有助于将其能力与其他大型语言模型进行比较。

排序理由 该条目报告了一个开源大型语言模型的基准测试结果,属于研究范畴。[lever_c_demoted from research: ic=1 ai=1.0]

在 Mastodon — fosstodon.org 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Sarvam 30B 模型性能指标在基准测试中公布

报道来源 [1]

  1. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    📊 Sarvam 30B (high) — 实际数字 GPQA: 63.3% 人类最后考试: 7.5% 长上下文推理: 0% SciCode: 19.2% 💰 每美元 136.2 智能点

    📊 Sarvam 30B (high) — the actual numbers GPQA: 63.3% Humanity's Last Exam: 7.5% Long Context Reasoning: 0% SciCode: 19.2% 💰 136.2 intelligence points per dollar Measured independently, not self-reported → https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI