PulseAugur
实时 09:25:57
English(EN) TradeVerse: A Longitudinal Benchmark of Political Negotiation in International Trade

新的TradeVerse基准测试用于评估LLM在国际贸易政治谈判中的能力

研究人员推出了TradeVerse,这是一个旨在评估大型语言模型(LLM)在理解国际贸易中纵向政治谈判能力的新基准测试。该基准测试由世界贸易组织记录的特定贸易关切构建而成,涵盖了1170次跨不同产品组会议的纪要。TradeVerse提出了三项任务:预测所讨论产品的协调制度(HS)编码,根据匿名会议内容识别响应国家,以及为谈判的最后一轮生成回应声明。该基准测试旨在突出这些复杂的多轮互动对当前LLM构成的挑战。 AI

影响 该基准测试有望推动LLM在国际贸易等专业领域处理复杂多轮对话能力的发展。

排序理由 该集群包含一篇介绍LLM新基准测试的研究论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的TradeVerse基准测试用于评估LLM在国际贸易政治谈判中的能力

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Debodeep Banerjee, Amitangshu Dasgupta ·

    TradeVerse:国际贸易中政治谈判的纵向基准

    arXiv:2608.06549v1 Announce Type: cross Abstract: LLMs are increasingly being applied to tasks involving institutional and political texts, but existing benchmarks evaluate them on isolated documents or single tasks. In realpolitik, negotiations are longitudinal data, where parti…