DCAS:解耦CLI智能体脚手架以在脚手架间内化规划能力
DCAS: Decoupling CLI Agent Scaffolding to Internalize Planning across Scaffolds
浏览论文内容
中文总结 AI 辅助
本文针对CLI智能体微调后跨脚手架性能下降问题,提出DCAS框架,通过解耦脚手架与后端实现跨脚手架评估,微调规划感知轨迹后模型在非训练脚手架下性能一致提升。
中文摘要 AI 辅助
基于命令行界面(CLI)的软件工程智能体已快速成熟,但开源生态系统已收敛至单一训练环境:用于微调开源模型的轨迹数据集几乎完全是在OpenHands平台下收集的。在该数据上微调的模型在OpenHands下表现良好,但部署到任何非训练脚手架时性能会大幅下降。未经过微调的基础模型不会出现这种差异,表明该差距由微调导致,且与训练脚手架的惯例相关。本文提出,脚手架特有的关键行为是规划结构,本文区分了两种规划形式:显式规划,即作为一等工件生成的执行前计划;隐式规划,即塑造智能体循环执行过程的结构惯例。基于该假设,缩小差距需要将规划从固定脚手架工件转变为学习到的模型能力。本文提出Decoupling CLI Agent Scaffolding(DCAS,解耦CLI智能体脚手架),这是一个后端替换拦截层,可在不修改脚手架的情况下路由任意CLI脚手架与任意后端模型之间的API流量,支持跨脚手架评估和规划感知轨迹收集。使用DCAS进行的受控计划源干预证实,规划质量是高杠杆率的组件,其收益超过了观察到的跨脚手架性能下降。在单个脚手架下,对一小部分DCAS收集的规划感知轨迹进行微调的模型,在非训练脚手架下表现一致提升,且两种规划形式在训练数据中可经验区分。
英文摘要
CLI-based software-engineering agents have matured rapidly, yet the open ecosystem has converged on a single training environment: trajectory datasets used to fine-tune open models are collected almost exclusively under OpenHands. Models fine-tuned on this data score well under OpenHands but degrade substantially when deployed under any non-training scaffold. Untrained base models do not show this divergence, indicating the gap is fine-tuning-induced and tied to the conventions of the training scaffold. We argue that a load-bearing scaffold-specific behavior is planning structure, in two senses this paper distinguishes: explicit planning, a pre-execution plan produced as a first-class artifact, and implicit planning, the structural conventions that shape execution throughout the agent loop. Under this hypothesis, closing the gap requires moving planning from a fixed scaffold artifact to a learned model capability. We introduce Decoupling CLI Agent Scaffolding (DCAS), a backend-substitution interception layer that routes API traffic between any CLI scaffold and any backend model without modifying the scaffold, enabling cross-scaffold evaluation and planning-aware trajectory collection. Using DCAS, a controlled plan-source intervention confirms planning quality is a high-leverage component, with gains exceeding the cross-scaffold drops we observe. A model fine-tuned on a small set of DCAS-collected planning-aware trajectories under a single scaffold gains consistently across non-training scaffolds, and the two senses of planning are empirically separable in training data.