编码智能体基准测试应匹配用户的任务流程
Coding-Agent Benchmarks Should Match Their Users' Task Flows
浏览论文内容
中文总结 AI 辅助
本研究通过分析真实软件工程师会话,发现编码智能体基准的任务流程与实际使用不符,提出SWE-TaskFlow方法转换基准以匹配目标任务流程,并强调交互协议是评估的重要维度。
中文摘要 AI 辅助
编码智能体的评估通常力求尽可能真实。在我们的研究中,我们收集了JetBrains IDE中真实软件工程师的4,782个智能体会话,称之为生产会话。由于我们的研究对象是交互式智能体,我们研究了至少包含三条用户消息的会话(占样本的33%)。这些长会话与基于问题的基准任务在两个方面有所不同:(i)用户请求涵盖了更广泛的任务类型混合——关于项目代码的问题、规划、审查、重构、执行——以及(ii)用户在整个会话中会在不同类型之间切换。来自三个公共交互语料库的长会话样本表现出显著不同的任务流程(会话长度、任务类型和类型间转换的分布),因此没有单一的交互分布是普遍真实的:基准测试应指定目标用例并校准到该用例的测量结果。我们提出了SWE-TaskFlow,一种转换任何基于问题的基准测试的方法:它保留已验证的任务和测试,同时通过提示拆分和可验证的仓库问答将交互引导至目标任务流程,并使用任务流程对齐分数(TFAS)在生成的轨迹中进行选择。在对700个SWE-Bench Pro任务的试点中,按顺序分几步解决任务大约使智能体成本翻倍,而解决率没有稳定变化:交互协议本身是评估的重要维度。
英文摘要
The evaluation of coding agents generally strives to be as realistic as possible. In our study, we collect 4,782 agent sessions of real software engineers in JetBrains IDEs, which we call Production Sessions. Since our subject is interactive agents, we study the sessions with at least three user messages (33% of the sample). These long sessions differ from issue-derived benchmark tasks in two ways: (i) user requests span a far wider mix of task types - questions about the project's code, planning, review, refactoring, execution - and (ii) users switch between types throughout a session. Long-session samples from three public interaction corpora exhibit markedly different Task Flows (the distributions of session lengths, task types, and type-to-type transitions), so no single interaction distribution is universally realistic: benchmarks should name a target use case and calibrate to measurements from it. We present SWE-TaskFlow, an approach for transforming any issue-derived benchmark: it preserves the verified tasks and tests while steering the interaction toward a target Task Flow through prompt splitting and verifiable repository QA, with a TaskFlow Alignment Score (TFAS) for selecting among generated trajectories. In a pilot on 700 SWE-Bench Pro tasks, solving the task sequentially in several steps approximately doubles agent cost without a stable change in resolve rate: the interaction protocol itself is an important dimension of evaluation.
发表机构
- JetBrains Research(JetBrains研究院)
机构由 AI 辅助整理,请以论文原文为准。