arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SKILL-KD:面向大语言模型智能体的对比式技能蒸馏

SKILL-KD: Contrastive Skill Distillation for LLM Agents

Qiming Shi, Yibo Dou, Jiawen Zhu, Yulong Tao, Linbo Jin, Zhaolu Kang, Yunfan Zhou, Di Weng

arXiv 2607.28048首次发表:更新:

AI 中文总结

SKILL-KD是面向LLM智能体的对比式技能蒸馏框架,通过将师生智能体的可操作差异蒸馏为技能补丁并迭代优化,结合感知漂移的技能整合,在五个智能体基准上提升了冻结学生智能体性能。

AI 中文摘要

基于技能的提示已成为改进大语言模型(LLM)智能体的实用机制,但现有技能获取方法常将技能视为经验总结、记忆条目或成功演示的直接总结,这给较弱的学生智能体带来了不匹配问题:当学生因缺乏任务知识或操作策略而失败时,其失败轨迹可能没有足够证据推断缺失的行为,而教师轨迹可能过于隐晦,难以内化为可复用的指导。我们提出SKILL-KD,这是一个对比式技能蒸馏框架,将技能视为不同能力智能体之间的显式蒸馏媒介。给定学生失败案例和同一任务的教师轨迹,SKILL-KD将二者的可操作差异蒸馏为文本式技能补丁,通过重新运行学生来评估该补丁,若学生仍失败则迭代优化补丁。为防止重复局部更新导致技能漂移,SKILL-KD还维护与轨迹关联的编辑历史,并执行感知漂移的技能整合,决定每个补丁是添加新规则、删除或修改现有规则,还是跳过。在五个智能体基准和两种学生设置下,SKILL-KD相较于固定模型适应基线,持续提升了冻结的学生智能体性能。

英文摘要

Skill-based prompting has become a practical mechanism for improving LLM agents, yet existing methods often treat skills as summaries of the agent's own experience or of successful demonstrations. This creates a mismatch for weaker student agents. A failed trajectory may not reveal the missing knowledge or strategy, while a teacher trajectory may be too implicit to internalize. We propose SKILL-KD, a contrastive skill distillation framework that treats skills as an explicit distillation medium between agents of different capabilities. Given a student failure and the teacher trajectory on the same task, SKILL-KD distills their actionable discrepancy into a textual skill patch, evaluates it by re-running the student, and iteratively refines it when the student still fails. To prevent skill drift from repeated local updates, SKILL-KD maintains trace-linked edit histories and performs Drift-Aware Skill Consolidation to decide whether each patch is added, merged, or skipped. Across five agent benchmarks and two student settings, SKILL-KD consistently improves frozen student agents over fixed-model adaptation baselines.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑