PulseAugur
EN
LIVE 06:29:32

Qwen3 Omni 30B A3B Instruct shows mixed benchmark results, excelling in some areas but failing in…

The Qwen3 Omni 30B A3B Instruct model has demonstrated strong performance on specific benchmarks, achieving 62% on GPQA and 72.5% on MMLU-Pro. However, it showed no capability in long-context reasoning. The model operates at 108.2 tokens per second and offers 14 "intel points" per dollar, according to independent measurements. AI

IMPACT This model's performance indicates areas of strength and weakness in current LLM capabilities, particularly highlighting the challenge of long-context reasoning.

RANK_REASON The item reports on specific benchmark results for an AI model, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]

Read on Mastodon — sigmoid.social →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Qwen3 Omni 30B A3B Instruct shows mixed benchmark results, excelling in some areas but failing in…

How we ranked this

Signal score
12 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
The item reports on specific benchmark results for an AI model, which falls under research. [lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
model release, paper
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

Full methodology in our editorial standards.

COVERAGE [1]

  1. Mastodon — sigmoid.social TIER_1 English(EN) · [email protected] ·

    📊 Qwen3 Omni 30B A3B Instruct: 62% GPQA, 72.5% MMLU-Pro, but 0% on long-context reasoning. At 108.2 tokens/sec and 14 intel points per dollar, independently mea

    📊 Qwen3 Omni 30B A3B Instruct: 62% GPQA, 72.5% MMLU-Pro, but 0% on long-context reasoning. At 108.2 tokens/sec and 14 intel points per dollar, independently measured → see the raw data. https:// olud.ai/leaderboard.html # LLM # Benchmarks # OpenSource # AI