PulseAugur
EN
LIVE 15:33:02

Azure GPT-5.6 variants fail invoice math, generating incorrect figures

A recent analysis of the azure/gpt-5.6-sol@low model variants reveals a significant flaw in their numerical reasoning capabilities. When presented with an invoice, these models consistently miscalculate subtotals and tax amounts, not by misreading digits, but by applying a flawed formula. Specifically, they appear to divide the total amount by 1.1, a calculation that yields an incorrect subtotal and tax value, even when the correct figures are clearly visible and the document's tax rate is not a simple 10%. This suggests the models are not performing arithmetic operations on visible data but are instead generating answers based on a predetermined incorrect formula. AI

IMPACT Reveals potential for LLMs to hallucinate numerical answers based on flawed internal logic rather than visible data.

RANK_REASON Analysis of model behavior on a specific task, revealing flaws. [lever_c_demoted from research: ic=1 ai=1.0]

Read on dev.to — LLM tag →

AI-generated summary · Google Gemini · from 1 sources. How we write summaries →

Azure GPT-5.6 variants fail invoice math, generating incorrect figures

COVERAGE [1]

  1. dev.to — LLM tag TIER_1 English(EN) · Hideki Mori ·

    The cheap tier reads one number

    <p>Here is a subtotal from a Japanese invoice rendered at 300 dpi, and the answer <code>azure/gpt-5.6-sol@low</code> gave for it. The printed value is ¥1,237,500, in the same body type this series has been testing all along. The answer was ¥1,235,000 — and it was ¥1,235,000 on <s…