Researchers have developed a new framework called SHAPE to analyze the Chain-of-Thought (CoT) reasoning processes of large language models (LLMs) in mathematical tasks. SHAPE examines how models interpret problems semantically and the specific mathematical heuristics they employ. The framework found that the heuristics used by a model are more indicative of correct answers than general CoT features, and that concentrating reasoning within fewer semantic spaces leads to better results, mirroring human behavior. Additionally, SHAPE revealed that reinforcement learning encourages specific heuristic usage, and a post-training method promoting diverse heuristics improved LLM accuracy. AI
IMPACT Provides a new diagnostic tool for understanding and improving LLM mathematical capabilities.
RANK_REASON The item is a research paper detailing a new framework for analyzing LLM reasoning. [lever_c_demoted from research: ic=1 ai=1.0]
- arXiv
- Chain-of-Thought
- Hugging Face
- Large language models
- mathematical reasoning
- reinforcement learning
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →