A new AI model, GPT-5.6 Sol, has achieved a significant milestone by reaching the ZeroBench human baseline at a pass@5 rate. This means that out of five attempts, at least one was correct, indicating a strong performance in complex problem-solving tasks. The model accomplished this without the use of external tools, highlighting its inherent reasoning capabilities. AI
IMPACT Demonstrates advanced reasoning capabilities, potentially setting new benchmarks for AI performance in complex tasks.
RANK_REASON Frontier-lab model release with benchmark result. [lever_c_demoted from frontier_release: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →