发表机构
Microsoft; University of California San Diego; Zhejiang University; Xi’an Jiaotong University; Nanjing University(微软; 加利福尼亚大学圣迭戈分校; 浙江大学; 西安交通大学; 南京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
Change2Task系统基于仓库历史将已合并拉取请求转为编码智能体任务,在5类任务中实现79.6%验证成功率,比基准多恢复29.2%任务,可减少流程支出10.8%。
AI 中文摘要
扩展编码智能体的规模需要持续提供可执行数据,用于训练、基准测试和持续评估,每个任务必须结合真实软件状态、规范、开发工具和可靠验证。为扩大这类数据供给,我们提出Change2Task系统,该系统基于仓库历史,将已合并的拉取请求转换为同一仓库健康现代修订版上的已验证任务。它将历史证据与演化代码对齐,通过补丁反转、代码映射或智能体重构来重建任务状态,并验证从健康基础到任务状态再到恢复状态的生命周期。通过从维护环境中基于开发者证据衍生多个任务,Change2Task为编码智能体的训练和评估提供可执行数据,同时减少重复的环境设置、存储和任务构建工作。我们通过5类常见且广泛采用的编码智能体任务家族(Bug修复、功能添加、测试生成、应用程序编程接口迁移、安全修复)对该系统进行评估。从1130个符合构建条件的源变更出发,Change2Task在这些任务家族中实现了79.6%的已验证任务构建成功率。在匹配的候选集中,它比基于拉取请求的构建基准多恢复29.2%的已验证任务。历史和重构案例在智能体评估下达到最高98.0%的匹配结果一致性,而现代基础的重复使用使整个流程的测量支出减少了10.8%。
英文摘要
Scaling coding agents requires a continuing supply of executable data for training, benchmarking, and continuous evaluation. Each task must couple a realistic software state with a specification, development tools, and reliable verification. To expand this supply, we present Change2Task, a system grounded in repository history that converts merged pull requests into verified tasks on healthy modern revisions of the same repository. It aligns historical evidence with evolved code, reconstructs task states through Patch Reversal, Code Mapping, or Agent Reconstruction, and validates the lifecycle from a healthy base to a task state and a restored state. By deriving multiple tasks grounded in developer evidence from maintained environments, Change2Task provides executable data for coding agent training and evaluation while reducing repeated environment setup, storage, and task construction effort. We evaluate the system through five common and widely adopted coding agent task families: Bug Fix, Feature Addition, Test Generation, Application Programming Interface Migration, and Security Repair. Starting from 1,130 source changes eligible for construction, Change2Task achieves 79.6% verified task construction success across these task families. On a matched candidate set, it recovers 29.2% more verified tasks than a construction baseline based on pull requests. Historical and reconstructed cases achieve up to 98.0% matched outcome agreement under agent evaluation, while reuse of modern bases reduces measured expenditure across the complete pipeline by 10.8%.
Comments15 pages, 7 figures, and 15 tables, including appendices