arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SWE-Touch:用户接触代码时的编码智能体基准测试

SWE-Touch: Benchmarking Coding Agents When Users Touch the Code

Yuqiao Tan, Jinxiang Meng, Fangyu Lei, Minzheng Wang, Shizhu He, Jun Zhao, Kang Liu

arXiv 2608.02499首次发表:更新:

发表机构

Institute of Automation, Chinese Academy of Sciences; University of Chinese Academy of Sciences(中国科学院自动化研究所; 中国科学院大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SWE-Touch基准测试显示,编码智能体对共享工作空间的状态感知不足,反编辑使其解决率降低7.7个百分点,需优化相关协作能力。

AI 中文摘要

实际软件开发要求编码智能体在共享工作空间中运行,用户可能在任务进行期间检查和修改代码,但现有的仓库级基准通常仅评估智能体单独工作,或将用户参与限制为消息交互。这促使我们提出问题:编码智能体如何理解并响应共享工作空间中的代码变更?我们引入SWE-Touch,这一框架通过经过验证的反编辑(Counter-Edits)对该场景进行压力测试,反编辑是与任务完成相冲突的、对任务相关代码的合理修改。SWE-Touch从多个修复轨迹中挖掘任务关键区域,使用独立的用户补丁生成器构建这些编辑,并在智能体到达相关代码时将其与上下文用户消息一同注入。我们在SWE-bench Verified上评估了9种编码模型,还在SWE-Bench Pro和DeepSWE的长周期任务上开展了额外实验。反编辑使SWE-bench Verified上的平均解决率降低了7.7个百分点,且这种性能下降在两个长周期基准上也持续存在。轨迹分析将这些失败归因于智能体对不断变化的工作空间的感知有限:智能体可能保留冲突代码,或在未充分重新检查仓库、未通过针对性测试验证修订后代码的情况下替换代码。这些发现表明,当前强大的自主性能仍无法确保共享工作空间协作所需的状态感知和自适应行为,同时指出检测工作空间变更、协调与任务冲突的编辑、验证受影响行为是未来优化的关键能力。

英文摘要

Real-world software development requires coding agents to operate in shared workspaces where users may inspect and modify code during an ongoing task, yet existing repository-level benchmarks typically evaluate agents working alone or restrict user participation to messages. This leads us to ask: how do coding agents understand and respond to code changes in a shared workspace? We introduce SWE-Touch, a framework that stress-tests this setting through validated Counter-Edits: plausible edits to task-relevant code that conflict with task completion. SWE-Touch mines task-critical regions from multiple repair trajectories, uses a separate User Patch Generator to construct the edits, and injects them with contextual user messages when agents reach the relevant code. We evaluate nine coding models on SWE-bench Verified, with additional experiments on longer-horizon tasks from SWE-Bench Pro and DeepSWE. Counter-Edit lowers average resolve rate by 7.7 percentage points on SWE-bench Verified, with degradation also persisting on both longer-horizon benchmarks. Trajectory analysis links these failures to limited awareness of the evolving workspace: agents may retain conflicting code or replace it without sufficiently re-inspecting the repository and validating the revised code with targeted tests. These findings show that strong autonomous performance does not yet ensure the state awareness and adaptive behavior needed for shared-workspace collaboration, and point to detecting workspace changes, reconciling conflicting edits with the task, and verifying the affected behavior as key capabilities for future optimization.

CommentsPreprint. Our code is available at https://github.com/Trae1ounG/SWE-Touch

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑