arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

XAgent:面向 GitHub 问题有效定位与解决的执行引导式智能体人工智能

XAgent: eXecution-guided Agentic AI for Effective Localization and Resolution of GitHub Issues

Hieu Huynh, Patanamon Thongtanunam, Michael Fu, Bach Le, Kla Tantithamthavorn

arXiv 2609.09769首次发表:更新:

发表机构

School of Computing and Information Systems, The University of Melbourne; Monash University(墨尔本大学计算机与信息学院; 蒙纳士大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

XAgent 通过执行引导分析动态行为与程序上下文,在 SWE-bench-lite 上实现 62.0% 解决率,优于现有方法,推动从静态描述向动态执行的 GitHub 问题解决范式转变。

AI 中文摘要

智能体人工智能已能够利用大型语言模型(LLM)自主解决仓库级别的 GitHub 问题。然而,由于依赖有限的问题静态描述,现有的智能体方法存在定位错误和验证不完整的问题。仅依赖此类信息可能使 LLM 的推理偏向于问题描述的狭窄范围,导致生成的补丁不完整,无法解决根本问题。在本文中,我们提出了 XAgent,一个执行引导的智能体框架,通过分析动态行为和额外的程序上下文来定位和验证问题。在 SWE-bench-lite 数据集上的实验结果表明,XAgent 优于其他现有方法,实现了 62.0% 的解决率和 72.8% 的函数定位准确率,同时保持了成本效率。我们的进一步分析表明,XAgent 成功解决了现有顶级基线未能解决的 7 个额外问题。这项工作强调了从静态的、面向描述的补丁生成向动态执行引导的问题解决的转变,为基于 LLM 的编码智能体实现更稳健和更通用的软件维护开辟了新的机遇。

英文摘要

Agentic AI has enabled capabilities in leveraging Large Language Models (LLMs) to autonomously resolve repository-level GitHub issues. However, due to the reliance on limited static description of issues, existing agentic approaches suffer from incorrect localization and incomplete validation. Solely relying on this information can bias LLM reasoning toward the narrow scope of the issue description, leading to incomplete patches that fail to address the underlying issue. In this paper, we present XAgent, an execution-guided agentic framework that analyzes dynamic behavior and additional program context to localize and validate issues. The experimental results on the SWE-bench-lite dataset demonstrate that XAgent outperforms other existing approaches, achieving a resolve rate of 62.0% and a function localization accuracy of 72.8%, while maintaining cost efficiency. Our analysis further shows that XAgent successfully resolves 7 additional issues that the top existing baselines fail to address. This work highlights a shift from static, description-oriented patch generation toward dynamic execution-guided issue resolution, opening new opportunities for LLM-based coding agents to achieve more robust and generalizable software maintenance.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑