arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

回到未来:用于电子表格创建基准测试的工作簿时光机

Back to the Future: A workbook time machine for spread sheet creation benchmarks

Mansi Uniyal, Agamdeep Singh, Ananya Singha, Priyanshu Gupta, Mukul Singh, Gust Verbruggen, Vu Le, Sumit Gulwani

arXiv 2608.07873首次发表:更新:

发表机构

Microsoft(微软公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出workbook time machine流水线,构建了含150个任务的电子表格创建评估基准wtmbench,揭示了查询特异性等因素对大语言模型Excel任务性能的关键影响。

AI 中文摘要

我们推出了workbook time machine(工作簿时光机),这是一种自动创建基准测试的流水线,用于评估语言模型创建电子表格中衍生对象的能力,包括公式、图表、数据透视表和条件格式。将其应用于公开工作簿语料库后,它生成了wtmcorpus——一组包含(输入工作簿、输出工作簿、查询)的三元组集合,涵盖四种对象类型且复杂度各异。我们从该语料库中整理出wtmbench,这是一个包含150个任务的评估基准,其查询分为三个特异性级别。我们在wtmbench上针对现有电子表格操作智能体和基线方法,从对象类型、步骤复杂度和指令粒度等维度进行评估。评估结果表明,查询特异性、智能体编排以及用于控制电子表格的接口API,对大语言模型在Excel任务上的性能有重大影响。

英文摘要

We introduce the workbook time machine, a pipeline that automatically creates benchmarks evaluating the ability of language models to create derived objects in spreadsheets (formulas, charts, pivot tables, and conditional formatting). Applied to public workbook corpora, it produces wtmcorpus--a collection of (input workbook, output workbook, query) triples spanning four artifact types and varying complexity. From this corpus we curate wtmbench, a 150-task evaluation benchmark with queries at three levels of specificity. We evaluate existing spreadsheet manipulation agents and baselines on wtmbench across artifact types, step complexity, and instruction granularity. Our evaluations show that query specificity, agent orchestration, and interface API used to control spreadsheets play a big role in LLM performance on Excel tasks.

CommentsCOLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑