PulseAugur
中
实时 00:00:05
English(EN) MathShikkha: A Controlled Study of Answer-Only and Chain-of-Thought Supervision for Bangla Mathematical Reasoning in Small Language Models

LLM中的孟加拉语数学推理:思维链监督效果喜忧参半

一项名为MathShikkha的新研究,调查了思维链(CoT)监督对于提高在孟加拉语上训练的小型语言模型数学推理能力的有效性。该研究使用GPT-5.4生成的推理过程构建了一个孟加拉语数学推理数据集,并对四种参数量从4B到7B不等的模型进行了微调。结果表明,CoT监督在较大的BanglaMATH基准测试中带来了显著的好处,所有模型的性能提高了20.1-28.1个百分点,而仅答案的微调有时会降低性能。然而,一项人类研究发现CoT在推理有效性方面没有显著提高,这表明其主要好处在于遵守目标语言和生成可检查的推理过程。 AI

影响 研究了不同监督方法对低资源语言数学推理的影响,为模型鲁棒性和可解释性提供了见解。

排序理由 学术论文,详细介绍了关于LLM推理能力的对照研究。[lever_c_demoted from research: ic=1 ai=1.0]

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

LLM中的孟加拉语数学推理:思维链监督效果喜忧参半

本文如何被排名

Signal score
0 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
学术论文,详细介绍了关于LLM推理能力的对照研究。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, model release
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
61 days old
Aged out of breaking-news scoring windows; ranking reflects the durable signal from the full source set.

完整方法见我们的编辑标准。

报道来源 [1]

  1. Hugging Face Daily Papers TIER_1 English(EN) ·

    MathShikkha:一项关于小型语言模型中孟加拉语数学推理的仅回答和思维链监督的对照研究

    Mathematical reasoning remains challenging in low-resource languages such as Bangla. We study whether teacher-generated Bangla Chain-of-Thought (CoT) supervision provides benefits beyond ordinary supervised fine-tuning. We construct \textsc{MathShikkha}, a Bangla mathematical rea…