PulseAugur
实时 09:17:52
English(EN) ComboShoppingBench: Evaluating LLM Agents for Budget-Constrained Basket Shopping with Coupons

新的基准 ComboShoppingBench 评估 LLM 代理执行复杂购物任务的能力

研究人员推出了 ComboShoppingBench,这是一个旨在评估大型语言模型 (LLM) 代理在复杂、预算受限的购物场景中的能力的新基准。该基准模拟了真实的组合购物任务,要求代理考虑商品兼容性、可用性、商店政策、配送费用、优惠券和预算限制。实验表明,即使是先进的 LLM 代理在处理这些任务时也面临困难,这表明在具有约束意识的购物能力方面仍有很大的改进空间。 AI

影响 凸显了当前 LLM 代理在预算受限购物等复杂现实世界决策任务中的局限性。

排序理由 该集群包含一篇介绍 LLM 代理新评估基准的研究论文。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的基准 ComboShoppingBench 评估 LLM 代理执行复杂购物任务的能力

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Adrian Li, Kelong Mao, Yudong Guo, Heming Xia, Xinwei Yang, Lirui Luo, Jace Wong, Pu Yao, Sulong Xu, Simiu Gu ·

    ComboShoppingBench:评估预算受限且含优惠券的购物篮购物LLM代理

    arXiv:2608.09282v1 Announce Type: new Abstract: Real-world shopping often requires constructing a basket of complementary items rather than retrieving a single product. Such combo-shopping tasks arise in device setup, meal preparation, event planning, and group takeout ordering, …