PulseAugur
实时 07:29:34
English(EN) UTP-Bench: Uncertainty-aware Travel Planning Benchmark

新基准UTP-Bench测试LLM在旅行规划中的不确定性

研究人员推出了UTP-Bench,这是一个旨在评估大型语言模型在不确定条件下生成旅行行程的鲁棒性的新基准。与假设确定性环境的先前基准不同,UTP-Bench整合了来自印度504个城市的真实世界数据,包括经验性延迟分布和人群模式。它提出了三个新指标——缓冲充分性得分、人群感知计时得分和交通延迟吸收得分——来量化生成的计划在处理交通延迟和人群可变性方面的表现。使用GPT-5、Qwen3、Mistral和Phi-4等模型进行的实验显示,与人类制定的计划相比,在时间缓冲和延迟感知调度方面存在显著的性能差距。 AI

影响 该基准可以推动LLM在需要不确定性下鲁棒规划的现实世界应用中的能力改进。

排序理由 该集群描述了一篇介绍用于评估LLM的基准的新学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新基准UTP-Bench测试LLM在旅行规划中的不确定性

本文如何被排名

Signal score
22 / 100
Composite score across the factors below. Higher = stronger signal that this story matters right now.
Newsworthiness bucket
Tool
该集群描述了一篇介绍用于评估LLM的基准的新学术论文。[lever_c_demoted from research: ic=1 ai=1.0]
Source corroboration
Single-source cluster
Only one publisher covered this so far. Single-source stories can still rank when the publisher is high-authority, but they lack cross-source corroboration.
Topics
paper, other
Editorial topic classification. Feeds into how the story surfaces on /topic/<slug> hub pages and into the per-entity coverage mix.
AI-industry relevance
High
Clearly on-topic for AI-industry coverage.
Story freshness
Breaking (< 6h)
Fresh story with cross-source coverage still developing. Ranking may shift as more sources report.

完整方法见我们的编辑标准

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Etcharla Revanth Rao, Priyanshu Karmakar, Shubhojit Mallick, Manish Gupta, Shreya Ghosh, Abhik Jana ·

    UTP-Bench:不确定性感知旅行规划基准

    arXiv:2609.02421v1 Announce Type: new Abstract: Large Language Models (LLMs) have recently demonstrated strong capabilities in automated travel itinerary generation. However, real- world travel planning is inherently uncertain: transportation delays, crowd fluctuations, and unexp…