PulseAugur
实时 12:49:01
English(EN) DS@GT ARC at Touch\'e: Large Language Models for Retrieval-Augmented Debate

研究发现,大型语言模型在辩论生成方面表现出色,但在评估方面存在困难

一篇研究论文详细介绍了 DS@GT ARCTouché 2025 检索增强辩论任务中的提交内容,该任务涉及使用来自三个供应商的六个领先的大型语言模型 (LLM)。该任务包括生成辩论发言并根据会话准则对其进行评估。研究发现,虽然前沿 LLM 在生成响应方面非常有效,并且在充当评估者时在其模型家族内部表现出高度一致性,但这种内部共识并不能可靠地预测官方评估性能,尤其是在“质量”准则方面。 AI

影响 强调了 LLM 评估能力的局限性,表明当前模型可能无法可靠地评估复杂任务中响应质量等细微方面。

排序理由 在 arXiv 上发表的研究论文,详细介绍了 LLM 在辩论任务中的表现。

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 2 个来源。 我们如何撰写摘要 →

研究发现,大型语言模型在辩论生成方面表现出色,但在评估方面存在困难

报道来源 [2]

  1. arXiv cs.CL TIER_1 English(EN) · Anthony Miyaguchi, Conor Johnston ·

    DS@GT ARC在Touch\'e上:用于检索增强辩论的大型语言模型

    arXiv:2608.08143v1 Announce Type: cross Abstract: We extend the DS@GT ARC working-note submission to the Touch\'e 2025 Retrieval-Augmented Debate task. The task has two subtasks: generating the next utterance in a simulated debate, and evaluating debate responses according to the…

  2. arXiv cs.IR (Information Retrieval) TIER_1 English(EN) · Conor Johnston ·

    DS@GT ARC 在 Touché:用于检索增强辩论的大型语言模型

    We extend the DS@GT ARC working-note submission to the Touché 2025 Retrieval-Augmented Debate task. The task has two subtasks: generating the next utterance in a simulated debate, and evaluating debate responses according to the Gricean maxims of Quantity, Quality, Relation, and …