PulseAugur
实时 17:47:26
English(EN) Humans Disengage, Reasoning Models Persist: Separating Difficulty Registration from Deliberation Allocation

研究发现,大型语言模型和人类在解决问题策略上存在分歧 · 已追踪 7 个来源

新研究表明,尽管人类和大型语言模型(LLMs)都会根据问题的难度调整解决时间,但其内部机制却存在显著差异。人类倾向于放弃那些他们认为困难或可能出错的问题,而大型语言模型则会在更难的问题上花费更多的计算资源,但这常常导致错误。这种“审议分配”上的分歧表明,大型语言模型在困难任务上延长处理时间源于不确定性,而非像人类那样进行战略性投入。 AI

影响 强调了大型语言模型和人类在处理复杂问题方式上的一个关键区别,表明当前大型语言模型的推理策略可能与人类的灵活性不完全一致。

排序理由 多篇 arXiv 论文讨论大型语言模型的推理能力和局限性。

在 Hugging Face Daily Papers 阅读 →

AI 生成摘要 · Google Gemini · 来自 15 个来源。 我们如何撰写摘要 →

研究发现,大型语言模型和人类在解决问题策略上存在分歧 · 已追踪 7 个来源

报道来源 [15]

  1. arXiv cs.AI TIER_1 English(EN) · Chenguang Wang, Ming Li, Xinyue Zeng, Zhuochun Li, Hong Jiao, Tianyi Zhou, Dawei Zhou ·

    LLM推理痕迹中的认知片段实现可解释的人工项目难度预测

    arXiv:2606.28186v1 Announce Type: cross Abstract: Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test construction. Existing methods often depend on costly human calibration or item-level textual repr…

  2. arXiv cs.AI TIER_1 English(EN) · Dawei Zhou ·

    LLM推理痕迹中的认知片段实现可解释的人工项目难度预测

    Predicting human item difficulty is central to educational assessment, where reliable estimates support fairness and effective test construction. Existing methods often depend on costly human calibration or item-level textual representations, providing limited evidence about the …

  3. arXiv cs.CL TIER_1 English(EN) · Bella Fascendini, Kathryn McGregor, Max D. Gupta, Thomas L. Griffiths ·

    谜题之谜:测试大型语言模型和人类的灵活推理能力

    arXiv:2606.27103v1 Announce Type: new Abstract: Humans flexibly adapt their reasoning strategies to the requirements of a given problem. Large language models (LLMs) have performed well on many cognitive tasks, however, it is unclear whether this accuracy is a result of pattern m…

  4. arXiv cs.CL TIER_1 English(EN) · Guan-Yi Lin, Hen-Hsen Huang ·

    大型模型优势何在:约束引导推理的优先性

    arXiv:2606.26108v1 Announce Type: new Abstract: Larger language models consistently outperform smaller ones on reasoning benchmarks, yet the reasoning differences underlying this gap remain underexplored. Across benchmarks in mathematics, physics, chemistry, and programming, we o…

  5. arXiv cs.AI TIER_1 English(EN) · Zhibo Hu, Chen Wang, Yanfeng Shu, Hye-young Paik, Liming Zhu ·

    隐喻是大语言推理模型跨领域失调的根源

    arXiv:2601.03388v3 Announce Type: replace-cross Abstract: Earlier research has shown that metaphors influence human decision-making, raising the question of whether metaphors also influence large language models (LLMs)' reasoning pathways, given that their training data contain a…

  6. arXiv cs.LG TIER_1 English(EN) · Hongyi Henry Jin, Wenhan Yang, Meysam Ghaffari, Carlos Morato, Baharan Mirzasoleiman ·

    推理能力早期显现:推理模型的数据策展

    arXiv:2606.26797v1 Announce Type: new Abstract: Supervised fine-tuning (SFT) on a small, high-quality set of long reasoning traces is an effective approach for eliciting strong reasoning capabilities in Large Language Models (LLMs). However, existing methods for curating high-qua…

  7. arXiv cs.AI TIER_1 English(EN) · Han-yu Wang ·

    人类退出,推理模型持续:区分难度注册与审议分配

    arXiv:2606.26502v1 Announce Type: new Abstract: Large reasoning models (LRMs) take longer on harder problems, just as humans do. This surface similarity hides an opposite pattern within items. When an LRM gets a problem wrong, it spends more tokens than when it gets the same prob…

  8. Hugging Face Daily Papers TIER_1 English(EN) ·

    LLM推理痕迹中的认知片段实现可解释的人工项目难度预测

    Epi2Diff framework transforms LRM reasoning traces into cognitive episodes to predict human item difficulty more accurately than existing methods.

  9. arXiv cs.CL TIER_1 English(EN) · Thomas L. Griffiths ·

    谜题的谜题:测试大型语言模型和人类的灵活推理能力

    Humans flexibly adapt their reasoning strategies to the requirements of a given problem. Large language models (LLMs) have performed well on many cognitive tasks, however, it is unclear whether this accuracy is a result of pattern matching from training data or flexible reasoning…

  10. arXiv cs.LG TIER_1 English(EN) · Baharan Mirzasoleiman ·

    推理质量早期显现:推理模型的数据策展

    Supervised fine-tuning (SFT) on a small, high-quality set of long reasoning traces is an effective approach for eliciting strong reasoning capabilities in Large Language Models (LLMs). However, existing methods for curating high-quality SFT data rely heavily on strong reasoning m…

  11. arXiv cs.CL TIER_1 English(EN) · Nicolas Boizard, Hippolyte Gisserot-Boukhlef, Kevin El Haddad, C\'eline Hudelot, Pierre Colombo ·

    规模还是推理?推理蒸馏的计算等效性分析

    arXiv:2509.22193v2 Announce Type: replace Abstract: Distilling reasoning traces from strong teacher models has become the standard recipe for building capable small language models. Yet reasoning traces are 5-20$\times$ longer than standard instruction fine-tuning (IFT) outputs, …

  12. Hugging Face Daily Papers TIER_1 English(EN) ·

    人类退出,推理模型持续:区分难度登记与审议分配

    Large reasoning models (LRMs) take longer on harder problems, just as humans do. This surface similarity hides an opposite pattern within items. When an LRM gets a problem wrong, it spends more tokens than when it gets the same problem right; humans do the reverse, spending less …

  13. arXiv cs.CL TIER_1 English(EN) · Han-yu Wang ·

    人类退出,推理模型持续:区分难度注册与审议分配

    Large reasoning models (LRMs) take longer on harder problems, just as humans do. This surface similarity hides an opposite pattern within items. When an LRM gets a problem wrong, it spends more tokens than when it gets the same problem right; humans do the reverse, spending less …

  14. dev.to — LLM tag TIER_1 English(EN) · Dhruv Aggarwal ·

    更快思考 vs. 更长思考:测试时计算

    <p>Imagine you are asked to solve a complex math problem. If you answer immediately, you’ll likely rely on intuition or a guess. But if you are given ten minutes to scratch out ideas, double-check your logic, and correct your mistakes before speaking, your accuracy improves signi…

  15. Mastodon — fosstodon.org TIER_1 English(EN) · [email protected] ·

    何时应让一个可能无法解决问题的推理模型放弃?Conformal Thinking 为测试时计算设定了无分布的停止/继续阈值

    When should a reasoning model quit a problem it probably can't solve? Conformal Thinking sets stop/continue thresholds for test-time compute by distribution-free risk control, holding the error rate under a target you pick. It adds a lower threshold that gives up early when confi…