A new arXiv paper explores the cost-effectiveness of using smaller, open-weight language models for grading mathematical proofs. The study found that models like GPT-OSS 120B, DeepSeek-V4 Flash, and Gemma-4 31B can achieve grading accuracy comparable to more expensive frontier models such as Claude Opus 4.7 and Gemini 3.1 Pro. The researchers suggest that requiring unanimous agreement among three budget models offers the highest accuracy and precision, though this specific rule requires independent replication. AI
IMPACT Demonstrates that cost-effective open-weight models can serve as viable alternatives for complex evaluation tasks, potentially lowering research costs.
RANK_REASON The cluster contains an academic paper detailing a new methodology and findings in AI research. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Claude Opus 4.7
- DeepSeek-V4 Flash
- Gemini-3.1 Pro
- Gemma 4.31B
- GPT-OSS 120B
- Hugging Face
- IMO-GradingBench
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →