发表机构
Peking University(北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对办公编辑中更新与保护并存的难题,提出170任务基准OfficeEditBench,区分交付与接受,揭示局部正确性不足,为保留逻辑的变更维护提供测试平台。
AI 中文摘要
一次小的办公编辑会产生两个义务:传播所有必需的更新,并保持受保护状态不变。更新过少会导致依赖关系不一致;更新过多则会改变用户未授权的内容。我们引入了OfficeEditBench,这是一个包含170个任务的基准,用于对电子表格、演示文稿和文档进行变更范围的维护。任务契约规定了必需的更新、受保护状态、原生结构和适用的交互要求。在来自WorkBuddy、Doubao和Codex的510个存档任务-系统结果中,我们区分了文件交付、目标完成和验证器定义的接受。硬包有效交付率从92%到100%不等,但没有选定的输出满足完整的契约。案例分析强调了为什么局部正确性是不够的:更新的值可能丢失其生成公式,修订的规则可能无法到达相关结论,新的截止日期可能省略保留的先决条件。这些机制将工件级检查与Office文件的持续可维护性联系起来。我们分析了维护失败,同时区分了冻结的自动判定与人类可接受性。OfficeEditBench提供了一个测试平台,用于在保留现有工作的逻辑和范围的同时完成所需的更改。
英文摘要
A small Office edit creates two obligations: propagate every required update and leave protected state untouched. Updating too little leaves dependencies inconsistent; updating too much changes content the user did not authorize. We introduce OfficeEditBench, a 170-task benchmark for change-scoped maintenance of spreadsheets, presentations, and documents. Task contracts specify required updates, protected state, native structures, and applicable interaction requirements. Across 510 archived task-system outcomes from WorkBuddy, Doubao, and Codex, we distinguish file delivery, target completion, and verifier-defined acceptance. Hard package-valid delivery ranges from 92% to 100%, yet no selected output satisfies the complete contract. Case analysis highlights why local correctness is insufficient: an updated value can lose its generating formula, a revised rule can fail to reach related conclusions, and a new deadline can omit a retained prerequisite. These mechanisms connect artifact-level checks to the continued maintainability of Office files. We analyze maintenance failures while distinguishing frozen automatic verdicts from human acceptability. OfficeEditBench provides a testbed for completing required changes while preserving the logic and scope of existing work.
Comments23 pages, 7 figures. Benchmark and code: https://github.com/Aniriswu/OfficeEditBench