Axiom, a seven-month-old startup, has achieved a significant milestone by solving 12 problems on the prestigious Putnam undergraduate math exam, scoring 8/12. This accomplishment places their AI system closer to top human performance than other reported AI systems. Axiom's CEO, Carina Hong, emphasizes that while coding ability is advancing rapidly, formal verification through tools like their open-sourced AXLE Lean toolkit is crucial for scaling and compounding AI brilliance, moving beyond informal proofs and statistical rewards. AI
IMPACT Demonstrates AI's growing capability in complex reasoning and formal proof generation, pushing the boundaries beyond current coding models.
RANK_REASON AI system achieves a notable benchmark result on a difficult human competition.
Read on Latent Space (podcast video) →
AI-generated summary · Google Gemini · from 2 sources. How we write summaries →