发表机构
IReSCoMath Research Laboratory, Faculty of Sciences, University of Gabes; National School of Engineering, Gabes(加贝斯大学理科学院 IReSCoMath 研究实验室; 加贝斯国立工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ASAD提出自适应智能体调试系统,根据缺陷复杂度动态配置智能体数量、角色与协作策略,在多个基准上较静态方法提升修复率12-20%,并减少32%智能体使用。
AI 中文摘要
将大型语言模型(LLMs)集成到多智能体系统中,已展现出在自动化调试方面的巨大潜力。然而,目前几乎所有的现有框架都依赖于僵化、预定义的架构:智能体的数量、角色及其交互模式在分析缺陷之前就已固定。这种“一刀切”的方法从根本上与软件缺陷的异构性不匹配。简单缺陷会因不必要的协调而浪费资源,而复杂缺陷则可能因专业知识不足或配置不当而受影响。本文介绍了ASAD,一个自适应的智能体调试系统,它根据每个缺陷的性质和复杂性来配置其团队。ASAD通过分析有缺陷的代码来启动调试过程,并动态确定要部署的智能体数量、它们应具备的专业角色以及应遵循的协作策略。一个中央协调器通过迭代规划、反思和执行来编排这一过程;对简单问题应用快速单次修复,同时组建专门团队来应对更复杂的故障。我们在三个既定基准上评估了ASAD:Defects4J、DebugBench和CodeFlaws,使用了多种LLM,如DeepSeek-V3、Qwen-3和GPT-5。ASAD在缺陷修复率上持续比思维链(CoT)提示提高12%至20%,在修复精度上持续优于静态多智能体系统4%至9%,同时将平均智能体使用量减少了32%。关键在于,我们的系统动态调整智能体的数量和角色:它以最少的协调解决简单缺陷,并且仅在更复杂的情况下扩展智能体的参与。
英文摘要
The integration of Large Language Models (LLMs) into multi-agent systems has shown great potential for automated debugging. Yet nearly all current frameworks rely on rigid, predefined architectures: the number of agents, their roles, and their interaction patterns are fixed before any analysis of the bug occurs. This one-size-fits-all approach is fundamentally mismatched to the heterogeneous nature of software defects. Simple bugs waste resources on unnecessary coordination, while complex ones suffer from insufficient or poorly aligned expertise. This paper introduces ASAD, an adaptive agentic system for debugging that configures its team according to the nature and complexity of each bug. ASAD initiates the debugging process by analyzing the faulty code and dynamically determines the number of agents to deploy, the specialized roles they should have, and the collaboration strategy they should follow. A central coordinator orchestrates this process through iterative planning, reflection, and execution; applying fast single-pass repairs for simple issues while assembling purpose-built teams to tackle more complex failures. We evaluate ASAD on three established benchmarks: Defects4J, DebugBench, and CodeFlaws, using multiple LLMs, such as DeepSeek-V3, Qwen-3 and GPT-5. ASAD consistently improves bug-fix rates by 12--20% over chain-of-thought(CoT) prompting and consistently outperforms static multi-agent systems by 4--9% in fix precision while reducing average agent usage by 32%. Crucially, our system dynamically adjusts the number and roles of agents: it resolves simple bugs with minimal coordination and scales agent involvement only for more complex cases.