arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.23377cs.SEcs.CLcs.LG

从一到多,从多到一:面向软件工程智能体的类别感知迭代专家训练

One to More, More to One: Category-Aware Iterative Expert Training for Software Engineering Agents

Jie Zhao, Ziyu Jiang, Suhang Zheng, Minghui Shan, Xiaoxiao Xu, Lin Qu

首次发表
浏览论文内容

中文总结 AI 辅助

针对仓库级软件工程任务在池化强化学习下的类别跷跷板问题,提出类别感知的专家训练与多教师在线策略蒸馏框架,在Pro-618和SWE-bench Multilingual上分别提升5.39和2.78个百分点。

中文摘要 AI 辅助

仓库级软件工程(SWE)包含异构的任务类别,在池化智能体强化学习下的进展可能是不均衡的:某些类别的提升伴随着其他类别的回退,而聚合的解决率掩盖了这些变化。受这种类别跷跷板现象的启发,我们开发了一个类别感知的专家训练与策略整合框架。可执行任务构建和SWE Labeler(一种基于证据的多轴标注系统)组织训练池。初始的类别特定强化学习提高了平均训练成功率,但留下了不均衡的实例级进展,这促使我们显式整合成功行为并进行策略自适应的任务选择。同源类别专家交替进行长视野的Agentic-miniRL与刷新-修复-扩展(RRE):更新后的策略刷新实例掌握度,重用其自身验证过的成功轨迹进行修复监督微调(Repair SFT),并重新选择任务进行进一步的强化学习。标签路由的多教师在线策略蒸馏(MOPD)将专家整合为一个可部署的学生模型,其中ReLU门控的奖励外推仅保留每个教师相对于参考模型的改进方向。专家训练和策略整合不需要外部模型提供解决方案轨迹或动作目标。我们通过聚合和每类别解决率、相对于每个联合强化学习基线的最小类别提升以及专家增益恢复来评估池化强化学习和平衡强化学习、专家发展以及单模型整合。最终的MOPD策略在Pro-618上达到58.04%的平均解决率,在SWE-bench Multilingual上达到59.00%,分别比基础模型提高了5.39和2.78个百分点。

英文摘要

Repository-level software engineering (SWE) comprises heterogeneous task categories, whose progress under pooled agentic reinforcement learning can be uneven: gains in some categories coincide with regressions in others, while aggregate resolution obscures these changes. Motivated by this category see-saw, we develop a category-aware expert-training and policy-integration framework. Executable task construction and SWE Labeler, an evidence-grounded multi-axis labeling system, organize the training pools. Initial category-specific RL improves average training success while leaving uneven instance-level progress, motivating explicit consolidation of successful behavior and policy-adaptive task selection. Same-origin category experts alternate long-horizon Agentic-miniRL with Refresh-Repair-Expand (RRE): the updated policy refreshes instance mastery, reuses its own verified successful trajectories for Repair SFT, and reselects tasks for further RL. Label-routed multi-teacher on-policy distillation (MOPD) consolidates the experts into one deployable student, with ReLU-gated reward extrapolation keeping only each teacher's improving direction over the reference. Expert training and policy integration require no external model to provide solution trajectories or action targets. We evaluate Pooled RL and Balanced RL, expert development, and single-model integration through aggregate and per-category resolution, the minimum category lift over each joint-RL baseline, and expert-gain recovery. The final MOPD policy achieves mean resolution of 58.04% on Pro-618 and 59.00% on SWE-bench Multilingual, improving over the base model by 5.39 and 2.78 percentage points, respectively.

发表机构

  • Alibaba Group(阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑