arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32600cs.AIcs.SE

CUA-SWE:当计算机使用智能体遇上可视化软件工程

CUA-SWE: When Computer-Use Agents Meet Visual Software Engineering

Prince Zizhuang Wang, Chenhao Liang, Zelong Xu, Aojie Yuan, Xiaolin Zhou, Haiyue Zhang, Yue Zhao, Xiyang Hu, Shuli Jiang

首次发表
浏览论文内容

中文总结 AI 辅助

CUA-SWE提出一个基准、环境和评估流程,研究计算机使用智能体在仅通过运行软件视觉界面获取信息时,如何结合源码执行与GUI交互完成软件工程任务,并通过确定性测试验证修复。

中文摘要 AI 辅助

软件开发需要的不仅仅是编辑代码:开发者需要反复运行软件、与其界面交互、目视检查其行为,并利用这些观察来决定接下来要更改什么以及某项更改是否有效。现有的编码智能体和计算机使用智能体在很大程度上是被孤立研究的,这使得这种集成开发过程尚未得到充分探索。诊断运行时交互故障要求智能体将视觉观察与负责的代码联系起来,然后再次使用该应用程序以验证修复是否成功。我们引入了CUA-SWE,一个用于计算机使用软件工程的基准、环境和评估流程。除了研究GUI反馈如何支持诊断和修复之外,我们还探讨了当所需规格或操作信息仅能通过运行中应用程序的视觉界面获得时,智能体是否能够完成软件工程任务。CUA-SWE涵盖四个软件工程领域,并要求智能体在同一任务中修改代码和配置、执行命令、与运行中的软件交互以及检查视觉反馈。每个任务都包含确定性的、特定于任务的测试,以验证最终软件是否满足要求并保留指定的行为。我们的评估刻画了前沿智能体如何将源代码级执行与应用程序截图和图形交互相结合,以产生经过验证的软件更改。我们考察了跨领域和任务信息需求的性能,以及成功修复相关的开发行为。CUA-SWE为研究智能体如何利用视觉反馈和交互来指导软件工程提供了一个统一的测试平台,并为最终软件提供了可执行的正确性标准。

英文摘要

Software development requires more than editing code: developers repeatedly run software, interact with its interfaces, visually inspect its behavior, and use these observations to decide what to change next and whether a change works. Existing coding agents and computer-use agents are largely studied in isolation, leaving this integrated development process underexplored. Diagnosing a runtime interaction failure requires agents to connect visual observations with the responsible code, then use the application again to verify the repair. We introduce CUA-SWE, a benchmark, environment, and evaluation pipeline for software engineering with computer use. Beyond studying how GUI feedback supports diagnosis and repair, we ask whether agents can complete software engineering tasks when required specification or operational information is available only through the running application's visual interface. CUA-SWE spans four software engineering domains and requires agents to modify code and configuration, execute commands, interact with running software, and inspect visual feedback within the same task. Each task includes deterministic, task-specific tests that verify whether the resulting software satisfies the requirements and preserves specified behavior. Our evaluation characterizes how frontier agents combine source-level execution with application screenshots and graphical interaction to produce verified software changes. We examine performance across domains and task information requirements, alongside the development behaviors associated with successful repairs. CUA-SWE provides a unified testbed for studying how agents use visual feedback and interaction to guide software engineering, with executable correctness criteria for the resulting software.

补充信息

↑