PulseAugur
EN
LIVE 22:05:52

Qwen3.8-27B achieves 29/30 on AIME 2026 math benchmark with FP8

A benchmark test of the Qwen3.8-27B model on the AIME 2026 math dataset revealed that its quantized FP8 weights, when set to xhigh reasoning effort, achieved a score of 29/30. This performance was comparable to the BF16 version at the same xhigh reasoning setting, but with significantly improved speed. The FP8 xhigh configuration also matched the BF16 medium setting's score while being faster. AI

IMPACT Demonstrates strong performance on complex reasoning tasks, potentially influencing future model development for mathematical problem-solving.

RANK_REASON Benchmark results for a specific model on a math dataset. [lever_c_demoted from research: ic=1 ai=1.0]

Read on r/LocalLLaMA →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Qwen3.8-27B achieves 29/30 on AIME 2026 math benchmark with FP8

COVERAGE [1]

  1. r/LocalLLaMA TIER_1 English(EN) · /u/No_Run8812 ·

    Qwen3.8-27B scored 29/30 on AIME 2026 with FP8 + xhigh reasoning — BF16 vs FP8 results

    <!-- SC_OFF --><div class="md"><p>I benchmarked Qwen3.8-27B on <code>MathArena/aime_2026</code> dataset, comparing BF16 and FP8 weights at medium and xhigh reasoning effort.</p> <h1>Interesting findings are:</h1> <ol> <li>quantized FP8 xhigh is better than BF 16 medium equally go…