Researchers have introduced HorizonMath, a new benchmark designed to assess AI's capabilities in solving complex, unsolved mathematical problems. The benchmark features 113 problems across eight domains, with a focus on the "generator-verifier gap" where discovery is difficult but verification is straightforward. Initial tests show that most current models score below 10%, but GPT-5.4 Pro and GPT-5.6 Sol each successfully discovered three novel solutions to research problems, contributing to the mathematical literature. AI
IMPACT Sets a new standard for evaluating AI's mathematical reasoning and discovery capabilities, potentially accelerating progress in AI-driven scientific research.
RANK_REASON The cluster is about a new academic paper introducing a benchmark and reporting research findings. [lever_c_demoted from research: ic=1 ai=1.0]
- alphaXiv
- arXiv
- CatalyzeX
- DagsHub
- Erik A. Wang
- Gotit.pub
- GPT 5.4-Pro
- GPT 5.6 "Sol"
- HorizonMath
- Hugging Face
- IArxiv
- ScienceCast
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →