A new study published on arXiv investigates how well large language models (LLMs) understand and generate mathematical problems that humans find interesting. Researchers compared LLM judgments of mathematical problem interestingness with those of crowdsourced participants and International Math Olympiad competitors. While LLMs generally align with human notions of interestingness, their judgment distributions and rationale correlations differ significantly from human preferences. The study also found that LLMs can generate valid and engaging math problems after filtering, suggesting potential for AI-human collaboration in mathematics. AI
IMPACT LLMs show potential as collaborators in mathematical reasoning, though their understanding of human interestingness requires further development.
RANK_REASON The cluster contains an academic paper detailing research findings on LLM capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →