PulseAugur
实时 08:43:00
English(EN) RetailBench: Benchmarking long horizon reasoning and coherent decision making of LLM agents in realistic retail environments

新的基准测试揭示大型语言模型(LLM)代理在复杂的购物和零售任务中存在困难

研究人员开发了新的基准测试来评估大型语言模型(LLM)代理在复杂、真实场景中的能力。ShoppingBenchEComAgentBench 专注于涉及隐藏意图、预算管理和多产品采购的复杂购物任务,揭示即使是 GPT-4.1 等先进模型也难以达到高成功率。同样,RetailBench 在长周期的零售管理模拟中评估 LLM 代理,突显了与最优策略相比,它们在决策和策略一致性方面存在显著差距。 AI

影响 这些基准测试强调了需要更强大的 LLM 代理来处理现实应用中复杂的多步推理和决策。

排序理由 多篇研究论文介绍了用于评估 LLM 代理在复杂、长周期任务中的新基准测试。

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 4 个来源。 我们如何撰写摘要 →

新的基准测试揭示大型语言模型(LLM)代理在复杂的购物和零售任务中存在困难

报道来源 [4]

  1. arXiv cs.CL TIER_1 English(EN) · Jiangyuan Wang, Kejun Xiao, Qi Sun, Huaipeng Zhao, Tao Luo, Jian Dong Zhang, Xiaoyi Zeng ·

    ShoppingBench:一个基于真实世界意图的购物基准测试,用于 LLM 驱动的智能体

    arXiv:2508.04266v4 Announce Type: replace Abstract: Existing benchmarks in e-commerce primarily focus on basic user intents, such as finding or purchasing products. However, real-world users often pursue more complex goals, such as applying vouchers, managing budgets, and finding…

  2. arXiv cs.AI TIER_1 English(EN) · Zeyao Du, Tong Li, Haibo Zhang ·

    EComAgentBench:在具有分布式隐藏意图的长期任务上对购物代理进行基准测试

    arXiv:2606.17698v1 Announce Type: new Abstract: As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly in the query, recorded in a profile, or revealed only when the right question is asked. Benchm…

  3. arXiv cs.CL TIER_1 English(EN) · Haibo Zhang ·

    EComAgentBench:在具有分布式隐藏意图的长期任务上对购物代理进行基准测试

    As LLM-based shopping agents enter production, existing benchmarks fail to capture how a shopper's requirements arrive: stated implicitly in the query, recorded in a profile, or revealed only when the right question is asked. Benchmarks that expose full intent upfront and grade o…

  4. arXiv cs.AI TIER_1 English(EN) · Linghua Zhang, Jun Wang, Jingtong Wu, Zhisong Zhang ·

    RetailBench:在真实的零售环境中对LLM代理的长远推理和连贯决策能力进行基准测试

    arXiv:2606.15862v1 Announce Type: new Abstract: Large language model (LLM) agents have made rapid progress on short-horizon, well-scoped tasks, yet their ability to sustain coherent decisions in dynamic long-horizon environments remains uncertain. We introduce RetailBench, a data…