PulseAugur
实时 09:22:35
English(EN) Back to the Future: A workbook time machine for spread sheet creation benchmarks

新的基准测试对LLM进行电子表格创建任务的测试

研究人员开发了一个“工作簿时光机”管道,以自动生成用于评估语言模型电子表格创建能力的基准。该管道生成输入-输出工作簿三元组的语料库和一个包含150个任务的精选基准wtmbench。对现有代理和基线的评估表明,查询特异性、代理编排和电子表格界面显著影响LLM在涉及Excel的任务上的表现。 AI

影响 该基准测试有望推动LLM在电子表格软件中执行复杂数据操作和生成任务的能力的提升。

排序理由 该集群描述了一篇介绍用于评估语言模型在电子表格任务上表现的基准的新学术论文。[lever_c_demoted from research: ic=1 ai=1.0]

在 arXiv cs.AI 阅读 →

AI 生成摘要 · Google Gemini · 来自 1 个来源。 我们如何撰写摘要 →

新的基准测试对LLM进行电子表格创建任务的测试

报道来源 [1]

  1. arXiv cs.AI TIER_1 English(EN) · Mansi Uniyal, Agamdeep Singh, Ananya Singha, Priyanshu Gupta, Mukul Singh, Gust Verbruggen, Vu Le, Sumit Gulwani ·

    回到未来:用于电子表格创建基准测试的“时间机器”工作簿

    arXiv:2608.07873v1 Announce Type: new Abstract: We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the ability of language models to create derived objects in spreadsheets (formulas, charts, pivot tables, and conditional formatting…