arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.29831cs.CLcs.SE

A^2Agent:面向仓库级代码定位智能体的动作感知强化学习方法

A^2Agent: Action-Aware Reinforcement Learning for Repository-Level Code Localization Agents

  • College of Computing and Informatics(计算与信息学院)
  • Sungkyunkwan University(成均馆大学)

机构由 AI 辅助整理,请以论文原文为准。

Doyeon Kim, Suyoung Bae, Yumin Lee, Jee-Hyong Lee

AI总结:

A^2Agent是一种动作感知强化学习方法,通过优化奖励与优势估计,在SWE-Bench的两个基准上提升代码定位性能,且4B模型优于更大规模的基线。

AI中文摘要:

定位与问题相关的代码区域是自动化软件工程的关键步骤。然而,现有方法依赖稀疏的轨迹级信号,无法识别每一轮动作的有效性,且常在探索过程中发现正确代码区域却无法提交。为解决这些局限,我们提出一种动作感知强化学习方法,结合每轮奖励序列(奖励发现与提交正确代码区域)与动作层面优势估计方案(通过分组共享相同探索上下文的轮次,隔离每个动作的贡献)。大量评估表明,该方法在SWE-Bench Verified上的平均F1值较现有最优方法(SOTA)提升1.58%,在SWE-Bench Pro上提升8.55%;我们的4B模型性能优于规模大至8倍的基线模型。代码可在指定URL获取。

英文摘要:

Localizing issue-relevant code regions is a critical step in automated software engineering. However, due to their reliance on sparse trajectory-level signals, existing methods cannot identify which per-turn actions are effective and often discover correct code regions during exploration but fail to commit them. To address these limitations, we propose an action-aware reinforcement learning method that combines a per-turn reward sequence rewarding both the discovery and commitment of gold code regions with an action-level advantage estimation scheme that isolates each action's credit by grouping turns sharing the same exploration context. Extensive evaluations show that our method improves the average F1 over the state-of-the-art (SOTA) by 1.58% on SWE-Bench Verified and 8.55% on SWE-Bench Pro, with our 4B model outperforming baselines up to 8x larger. Our code is available at https://github.com/donian00/A2Agent.

补充信息

↑