OpenAI has detailed how specific API settings significantly impact benchmark performance, particularly for their GPT-5.6 "Sol" model. By enabling "retained reasoning" and "context compaction" through the Responses API, GPT-5.6 Sol saw an 188% score increase on the ARC-AGI-3 benchmark while using six times fewer output tokens. OpenAI suggests these settings, along with proper harness design and prompting, are crucial for accurately evaluating models, especially for long-running agent tasks where memory and context management are key. AI
IMPACT Highlights the importance of standardized evaluation practices and API configurations for accurate AI model performance measurement.
RANK_REASON OpenAI is explaining how their own API settings affect benchmark results, rather than announcing a new model or capability.
AI-generated summary · Google Gemini · from 5 sources. How we write summaries →