A new study, MathShikkha, investigated the effectiveness of Chain-of-Thought (CoT) supervision for improving mathematical reasoning in small language models trained on the Bangla language. The research constructed a Bangla mathematical reasoning dataset using GPT-5.4-generated rationales and fine-tuned four models ranging from 4B to 7B parameters. Results indicated that CoT supervision provided significant benefits on the larger BanglaMATH benchmark, improving performance by 20.1-28.1 points across all models, while answer-only fine-tuning sometimes degraded performance. However, a human study found no significant improvement in reasoning validity from CoT, suggesting its primary benefits lie in adherence to the target language and producing inspectable reasoning. AI
IMPACT Investigates the impact of different supervision methods on low-resource language mathematical reasoning, offering insights into model robustness and interpretability.
RANK_REASON Academic paper detailing a controlled study on LLM reasoning capabilities. [lever_c_demoted from research: ic=1 ai=1.0]
Read on Hugging Face Daily Papers →
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →