An open-source AI research system named PRAXIST has achieved superior results on the MLE-bench compared to a Claude Code baseline. PRAXIST secured 49 gold medals for approximately $3,000 in compute costs, while the Claude Code baseline achieved 34 gold medals at a cost of $38,000. The system's success is attributed to its method of inheriting evidence from failed experiments rather than just scores, allowing lessons to be passed to subsequent generations. AI
IMPACT Demonstrates a novel approach to AI research that significantly reduces compute costs while improving performance, potentially influencing future AI development methodologies.
RANK_REASON The item describes a research system's performance on a benchmark, detailing its methodology and results. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Mastodon — mastodon.social →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →