arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StagedWorkspace:面向知识工作智能体的版本化工作空间

StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents

Yining Hua, Hongbin Na, Yifan Zhou, Akshay Kalose, Cyrus Ayubcha, Levi Lian

arXiv 2608.18050首次发表:更新:

发表机构

Harvard University; Raycaster AI; University of Technology Sydney; University of Washington; Stanford University(哈佛大学; 雷caster人工智能公司; 悉尼科技大学; 华盛顿大学; 斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对知识工作智能体的版本化工作空间问题,提出StagedWorkspace方案,通过绑定解析记录与原生文件内容哈希提升性能,在OfficeQA Pro等数据集上取得显著优于基准的效果,为相关基准构建提供了新方向。

AI 中文摘要

AI智能体越来越多地执行知识工作,即生成或修改代码仓库、文档、电子表格、幻灯片、报告等持久化数字工件,但它们所搜索的解析视图、所编辑的原生文件、所审核的变更以及所提交的工件,可能指向同一工作产品的不同版本。我们将此问题定义为工作空间状态契约:每个视图都应明确关联到不断演变的工作空间状态的某个版本。编码智能体通过针对搜索、差异对比(diffs)和测试的仓库契约部分满足了这一需求,而对于PDF、电子表格、幻灯片、笔记以及混合格式项目文件夹,类似的契约则不够明确。我们提出StagedWorkspace,一种面向知识工作智能体的版本化工作空间,该工作空间会在原生文件发生变更时,将解析记录和审核差异对比绑定到其内容哈希值上。在OfficeQA Pro和APEX-Agents上进行的固定框架消融实验显示,双解析/原生访问在所有测试模型中均取得最高点估计值;与更受限的单一视图相比,该方式将OfficeQA Pass@1提升了8.3至12.1个百分点,将APEX平均评分提升了4.7至9.2分。SW-AGENT在OfficeQA上使用Gemini 3.1 Pro的得分为63.9%,在APEX上使用GPT-5.4 Nano的得分为42.1,而已发表的相同模型得分分别为29.3%和25.5。在针对57项文件编辑任务的配对审核轴消融实验中,进一步发现当差异对比可见时,观测到的分数更高。这些结果表明,工作空间状态是知识工作智能体的一个实验变量,并催生了将证据、分阶段编辑和提交的工件作为显式状态转换进行评分的基准。

英文摘要

AI agents increasingly perform knowledge work (i.e., produce and modify persistent digital artifacts such as code repositories, documents, spreadsheets, slides, reports), yet the parsed views they search, the native files they edit, the changes they review, and the artifacts they submit can refer to different versions of the same work product. We formulate this as a workspace-state contract: every view should be explicitly tied to a version of the evolving workspace state. Coding agents partly address this need through repository contracts for search, diffs, and tests, whereas an analogous contract is less explicit for PDFs, spreadsheets, slides, notebooks, and mixed-format project folders. We propose StagedWorkspace, a versioned workspace for knowledge-work agents. The workspace binds parsed records and review diffs to content hashes of the native files as they change. In fixed-harness ablations on OfficeQA Pro and APEX-Agents, dual parsed/native access has the highest point estimate for every tested model; relative to the more limiting single view, it improves OfficeQA Pass@1 by 8.3-12.1 points and APEX mean rubric score by 4.7-9.2 points. SW-AGENT scores 63.9% with Gemini 3.1 Pro on OfficeQA and 42.1 with GPT-5.4 Nano on APEX, compared with published same-model scores of 29.3% and 25.5, respectively. A paired review-axis ablation on 57 file-editing tasks further finds higher observed scores when diffs are visible. These results identify workspace state as an experimental variable in knowledge-work agents and motivate benchmarks that score evidence, staged edits, and submitted artifacts as explicit state transitions.

CommentsUnder Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑