Cadence:面向编码智能体的策略性指导
Cadence: Strategic Guidance for Coding Agents
浏览论文内容
中文总结 AI 辅助
本文针对现有编码智能体监控器的静态触发方案局限,提出动态监控框架Cadence,其含双层干预模块与检查调度器,在SWE-bench Lite任务上显著提升编码智能体的问题解决率,且令牌效率具竞争力。
中文摘要 AI 辅助
运行时监控器正越来越多地用于提升基于大语言模型(LLM)的编码智能体的可靠性,其通过检查执行轨迹并在检测到异常行为时提供纠正性指导。然而,现有监控器的有效性受限于静态的指导触发方案:要么依赖固定的检查间隔,在严重异常行为期间无法及时提供指导,而在正常执行期间产生不必要的开销;要么依赖僵化的启发式规则,无法检测复杂的推理错误。为解决这些局限,本文提出Cadence,这是一种动态监控框架,可根据智能体的实时执行健康状况自适应地安排检查并提供指导。Cadence包含两个核心模块:双层干预模块和检查调度器。干预模块针对正常执行和轻微失误提供咨询级指导,针对严重异常行为提供替换级指导。调度器根据干预级别调整检查频率:在提供替换级指导后收紧监督,在提供咨询级指导后放松监督。在两个不同智能体(mini-swe-agent和Moatless)的300个SWE-bench Lite任务上进行评估,Cadence在所有评估的监控器中达到最高的解决率。具体而言,Cadence在mini-swe-agent上比普通智能体的解决率高出25.33%(多解决76个任务),在Moatless上高出15.67%(多解决47个任务),同时与最先进的基线相比保持了有竞争力的令牌效率。
英文摘要
Runtime monitors are increasingly used to improve the reliability of LLM-based coding agents by inspecting execution trajectories and delivering corrective guidance upon detecting misbehavior. However, their effectiveness remains limited by static guidance triggering schemes. Existing monitors rely either on fixed inspection intervals, missing timely guidance during severe misbehaviors while incurring unnecessary overhead during healthy execution, or on rigid heuristic rules, failing to detect complex reasoning errors. To address these limitations, we propose Cadence, a dynamic monitoring framework that adaptively schedules inspections and delivers guidance according to the agent's real-time execution health. Cadence consists of two core modules: a two-tier intervention module and an inspection scheduler. The intervention module delivers advisory-level guidance for normal executions and minor lapses, while providing replacement-level guidance for severe misbehaviors. Driven by the intervention level, the scheduler adjusts the inspection frequency by tightening supervision after replacement-level guidance and relaxing it after advisory-level guidance. Evaluated on 300 SWE-bench Lite tasks across two distinct agents, mini-swe-agent and Moatless, Cadence achieves the highest resolve rate among all evaluated monitors. Specifically, Cadence outperforms vanilla agents by 25.33\% (+76 resolved tasks) on mini-swe-agent and 15.67\% (+47 resolved tasks) on Moatless, while maintaining competitive token efficiency compared to state-of-the-art baselines.
发表机构
- Singapore Management University(新加坡管理大学)
- University College London(伦敦大学学院)
机构由 AI 辅助整理,请以论文原文为准。