PulseAugur
实时 07:21:21
English(EN) TREK: A Travel Reasoning and Evaluation Kit for LLM Agents in Complex Trip Planning

新的 TREK 基准测试 LLM 智能体在复杂行程规划方面的能力

研究人员推出了 TREK,这是一个旨在严格评估 LLM 智能体在复杂行程规划场景中能力的新基准。TREK 通过专注于生成完全可行、符合约束、无幻觉、时空可执行且预算有效的行程,同时响应未明确说明的旅行者需求,来解决现有基准的局限性。该基准包含 800 个多约束任务和一个确定性评估器,发现即使是 GPT-5.6 等先进模型,在可解决任务中也只能达到 46.2% 的完全可行性,而未明确说明的需求构成了重大挑战。 AI

影响 该基准将有助于研究人员识别并解决 LLM 智能体规划能力中的关键局限性,特别是在未明确说明的用户需求和整体可行性方面。

排序理由 该集群包含一篇学术论文,介绍了一个用于 LLM 智能体的新基准和评估工具包。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.CL 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的 TREK 基准测试 LLM 智能体在复杂行程规划方面的能力

报道来源 [1]

  1. arXiv cs.CL TIER_1 English(EN) · Jinhu Qi, Wentao Zhang, Siu Man Ng, Feiyang Xu, Yanyu Chen, Yaoman Li, Irwin King ·

    TREK: 用于复杂行程规划中 LLM 智能体的旅行推理与评估工具包

    arXiv:2607.26977v1 Announce Type: new Abstract: Travel planning is a demanding stress test for tool-using LLM agents: a usable itinerary is a single artifact that must be right along many axes at once - every flight, hotel, and attraction must exist and be bookable, the days must…