A recent analysis of the azure/gpt-5.6-sol@low model variants reveals a significant flaw in their numerical reasoning capabilities. When presented with an invoice, these models consistently miscalculate subtotals and tax amounts, not by misreading digits, but by applying a flawed formula. Specifically, they appear to divide the total amount by 1.1, a calculation that yields an incorrect subtotal and tax value, even when the correct figures are clearly visible and the document's tax rate is not a simple 10%. This suggests the models are not performing arithmetic operations on visible data but are instead generating answers based on a predetermined incorrect formula. AI
IMPACT Reveals potential for LLMs to hallucinate numerical answers based on flawed internal logic rather than visible data.
RANK_REASON Analysis of model behavior on a specific task, revealing flaws. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →