PulseAugur
实时 06:41:29
English(EN) 📊 Qwen3 Omni 30B A3B Instruct: 62% GPQA, 72.5% MMLU-Pro, but 0% on long-context reasoning. At 108.2 tokens/sec and 14 intel points per dollar, independently mea

Qwen3 Omni 30B A3B Instruct 的基准测试结果喜忧参半,在某些领域表现出色,但在长上下文推理方面失败

Qwen3 Omni 30B A3B Instruct 模型在特定基准测试中表现强劲,GPQA 达到 62%,MMLU-Pro 达到 72.5%。然而,它在长上下文推理方面没有表现出任何能力。根据独立测量,该模型运行速度为每秒 108.2 个 token,每美元提供 14 个“智能点”。 AI

影响 该模型的性能表明了当前 LLM 能力的优势和劣势,特别突出了长上下文推理的挑战。

排序理由 该项目报告了 AI 模型的具体基准测试结果,属于研究范畴。[lever_c_从研究降级:ic=1 ai=1.0]

在 Mastodon — sigmoid.social 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

Qwen3 Omni 30B A3B Instruct 的基准测试结果喜忧参半,在某些领域表现出色,但在长上下文推理方面失败

本文如何被排名

Signal score
11 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该项目报告了 AI 模型的具体基准测试结果,属于研究范畴。[lever_c_从研究降级:ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    📊 Qwen3 Omni 30B A3B Instruct:GPQA 62%,MMLU-Pro 72.5%,但在长上下文推理方面为 0%。速度为 108.2 tokens/秒,每美元 14 个智能点,独立测量

    📊 Qwen3 Omni 30B A3B Instruct: 62% GPQA, 72.5% MMLU-Pro, but 0% on long-context reasoning. At 108.2 tokens/sec and 14 intel points per dollar, independently measured → see the raw data. https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI