发表机构
Tencent Hy AI Data; Beijing Zhongguancun Academy(腾讯Hy AI数据; 北京中关村学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
SWE-Journey是用于更真实评估代码助手的基准,通过弱到强合成流水线构建长周期编码任务、用户模拟智能体复现真实交互,发现当前代码助手与非编程人员交互时功能测试通过率不足25%,关键能力为正确提问、查找与修复。
AI 中文摘要
Claude Code、Codex等代码助手已成为大语言模型智能体的重要应用,但现有基准与实际使用场景仍存在较大差距,尤其在任务周期和交互长度方面。代码助手需要在持续演进的代码仓库中完成长链开发工作,同时通过多轮交互反复明确需求并调整实现方案。为解决这些差距,我们提出SWE-Journey——一个用于更真实评估代码助手的基准。针对任务周期差距,我们提出弱到强合成流水线,自动构建长周期编码任务;针对交互差距,我们从真实交互数据中挖掘出四类代表性用户角色,并构建用户模拟智能体以复现真实的代码辅助交互。平均而言,模型在与软件架构师交互时能通过超过75%的功能测试,但与非编程人员交互时通过率不足25%。这些结果表明,当前代码助手仍无法为非编程人员提供可靠的编码支持。我们进一步分析了该差距的原因,确定了交互过程中“正确提问”“正确查找”和“正确修复”是关键能力。
英文摘要
Coding assistants such as Claude Code and Codex have become a major application of LLM agents, yet existing benchmarks remain far from real-world use, particularly in task horizon and interaction length. Code assistants require completing long chains of development work in continuously evolving repositories, while repeatedly clarifying requirements and adapting implementations through multi-turn interaction. To address these gaps, we introduce SWE-Journey, a benchmark for more realistic evaluation of coding assistants. To address the task-horizon gap, we propose a weak-to-strong synthesis pipeline that automatically constructs long-horizon coding tasks. To address the interaction gap, we mine four representative user personas from real interaction data and build a user-simulation agent to reproduce realistic code-assistance interactions. On average, models pass over 75% of tests for requested functionality with software architects, but fewer than 25% with non-coders. These results show that current coding assistants still fall short of enabling reliable coding for non-coders. We further analyze the reasons for this gap and identify asking right, finding right, and fixing right as key capabilities during interaction.