分析与缓解编码智能体中的成本低效行为
Analyzing and Mitigating Cost-Inefficient Behaviors in Coding Agents
浏览论文内容
中文总结 AI 辅助
本研究首次分析编码智能体成本低效行为,发现三种常见模式,并比较三种缓解策略,证明开发者设计技能可显著降低成本。
中文摘要 AI 辅助
尽管编码智能体有效,但它们往往会产生大量的金钱成本。它们反复出现的成本低效行为仍未得到充分探索。我们开展了对编码智能体中行为成本低效的首次研究,分析了来自Claude Code和Mini-SWE-Agent在SWE-bench Verified上四种配置下的1,200条轨迹。我们识别出三种成本低效行为:子sumed检索、相似脚本生成和测试重新执行。然后,我们在保留的SWE-bench Verified和Pro任务上评估了三种缓解策略:结构感知检索、智能体合成技能和开发者设计技能,共10k条轨迹。我们的主要发现是:(1)这三种行为影响了79.00%至98.00%的编码任务,并占任务成本的22.75%。(2)结构感知检索可能引入检索开销并改变智能体委托,导致检索效率改进不一致,成本增加高达28.14%。(3)智能体合成技能往往产生低级、轨迹特定的指导,限制了其有效性和通用性。(4)相比之下,开发者设计技能提供高级、轨迹无关的指导,将成本降低高达41.73%,大约是智能体合成技能最大增益的两倍。
英文摘要
Although effective, coding agents often incur substantial monetary costs. Their recurring cost-inefficient behaviors remain underexplored. We conduct the first study of behavioral cost inefficiencies in coding agents, analyzing 1,200 trajectories from Claude Code and Mini-SWE-Agent across four configurations on SWE-bench Verified. We identify three cost-inefficient behaviors: subsumed retrieval, similar script generation, and test re-execution. We then evaluate three mitigation strategies: structure-aware retrieval, agent-synthesized skills, and developer-designed skills, over 10k trajectories on held-out SWE-bench Verified and Pro tasks. Our main findings are: (1) The three behaviors affect 79.00%--98.00% of coding tasks and account for up to 22.75% of task cost. (2) Structure-aware retrieval can introduce retrieval overhead and alter agent delegation, causing inconsistent improvements in retrieval efficiency and cost increases of up to 28.14%. (3) Agent-synthesized skills tend to produce low-level, trace-specific guidance, limiting their effectiveness and generality. (4) In contrast, developer-designed skills provide high-level, trace-agnostic guidance, reducing cost by up to 41.73%, roughly twice the maximum gain from agent-synthesized skills.
发表机构
- Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。