Researchers have introduced MathNet, a new multimodal and multilingual dataset designed to evaluate the mathematical reasoning and retrieval capabilities of large language models. The dataset comprises over 30,000 Olympiad-level math problems from 47 countries and 17 languages, spanning two decades. Initial experiments show that current state-of-the-art models like Gemini-3.1 Pro and GPT-5 still struggle with these complex problems, while retrieval-augmented generation models, such as DeepSeek-V3.2-Speciale, demonstrate significant performance improvements. AI
IMPACT This benchmark could drive improvements in AI's mathematical reasoning and retrieval capabilities, particularly in multilingual contexts.
RANK_REASON The cluster describes a new academic paper introducing a benchmark dataset for AI research. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →