PulseAugur
实时 05:56:28
English(EN) When Contextual Inference Fails: Cancelability in Interactive Instruction Following

新基准测试大型语言模型推断上下文和澄清指令的能力

研究人员开发了一个名为 Build What I Mean (BWIM) 的新基准,用于测试大型语言模型在交互式指令遵循任务中处理上下文推理和可取消性的能力。该基准模拟了一个协作式积木搭建场景,模型必须在指令不明确时决定是进行推断还是寻求澄清。评估显示,尽管模型能够检测说话者的不可靠性,但它们难以相应地调整其行为,在不确定性下常常默认采用低效的澄清策略或进行猜测。 AI

影响 这项研究可能有助于开发出更强大的大型语言模型,使其能够在复杂任务中进行细致的交互并更好地理解人类意图。

排序理由 该集群是关于一篇详细介绍用于评估大型语言模型能力的新基准的学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新基准测试大型语言模型推断上下文和澄清指令的能力

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Natalia Bila, Kata Nasz\'adi, Alexandra Mayn, Christof Monz ·

    当上下文推理失败时:交互式指令遵循中的可取消性

    arXiv:2603.19997v2 Announce Type: replace Abstract: We investigate the separation of literal interpretation from contextual inference in a collaborative block-building tasks, where an agent must resolve underspecified instructions using context. We adapt an existing two-speaker p…