CodeSkill:面向长程代码智能体的潜在技能抽象
CodeSkill: Latent Skill Abstraction for Long-Horizon Code Agents
AI总结:
CodeSkill通过层次化潜在技能建模,将代码智能体的强化学习从令牌级提升至经验级,利用教师模型提炼轨迹并注入LLM,显著提升长程决策效率与跨域泛化能力。
AI中文摘要:
代码智能体需要在复杂的交互轨迹上进行长程决策。然而,现有的强化学习(RL)方法通常在令牌级别优化行为,导致低级生成与高级行为推理之间的不匹配。这一局限导致在稀疏奖励下的探索效率低下和信用分配薄弱。此外,尽管大规模智能体轨迹中常包含重复的多步行为模式,但其噪声化的令牌级表示阻碍了有效的经验复用。为应对这些挑战,我们提出了CodeSkill,一个将层次化潜在技能建模适配到代码智能体领域的框架。CodeSkill首先利用教师模型将成功和失败的轨迹提炼为多级文本抽象。然后,它将时间变分推断与强化学习相结合,将这些离散语义映射为连续潜在变量,同时自适应边界机制根据执行反馈动态门控技能转换。学习到的技能作为潜在语义前缀注入冻结的LLM策略中,从而在紧凑的语义空间而非原始令牌序列上进行优化。通过将RL从令牌级探索转变为经验级推理,CodeSkill提升了优化效率和长程行为连贯性。大量实验表明,CodeSkill在多样化的通用和工业代码基准上,相较于强开源权重基线取得了极具竞争力的性能。此外,学习到的技能展现出强大的可迁移性和稳健的跨领域泛化能力,凸显了显式行为抽象对可扩展智能体代码生成的有效性。
英文摘要:
Code agents require long-horizon decision-making over complex interaction trajectories. However, existing reinforcement learning (RL) approaches typically optimize behavior at the token level, creating a mismatch between low-level generation and high-level behavioral reasoning. This limitation leads to inefficient exploration and weak credit assignment under sparse rewards. Moreover, while large-scale agent trajectories often contain recurring multi-step behavioral patterns, their noisy token-level representations hinder effective experience reuse. To address these challenges, we propose CodeSkill, a framework that adapts hierarchical latent skill modeling to the code agent domain. CodeSkill first leverages a teacher model to distill both successful and failed trajectories into multi-level textual abstractions. It then integrates temporal variational inference with reinforcement learning to map these discrete semantics into continuous latent variables, while an adaptive boundary mechanism dynamically gates skill transitions based on execution feedback. The learned skills are injected into a frozen LLM policy as latent semantic prefixes, enabling optimization in a compact semantic space rather than over raw token sequences. By shifting RL from token-level exploration to experience-level reasoning, CodeSkill improves optimization efficiency and long-horizon behavioral coherence. Extensive experiments demonstrate that CodeSkill achieves highly competitive performance against strong open-weight baselines across diverse general and industrial coding benchmarks. Furthermore, the learned skills exhibit strong transferability and robust cross-domain generalization, highlighting the effectiveness of explicit behavioral abstraction for scalable agentic code generation.