arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SkillGate:训练长视距智能体中的策略内技能选择

SkillGate: Training In-Policy Skill Selection in Long-Horizon Agents

Qingyao Li, Wenxiang Jiao, Shuai Shao, Kangning Zhang, Yuan Lu, Yi Guo, Weiwen Liu, Weinan Zhang, Yong Yu

arXiv 2608.18852首次发表:更新:

发表机构

Shanghai Jiao Tong University; Xiaohongshu Inc.(上海交通大学; 小红书科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

SkillGate 针对长视距智能体策略内技能选择的选择器信用匮乏问题,通过划分信用通道提升了 9B 策略的试验成功率,减少了误导性候选接触并降低技能读取次数。

AI 中文摘要

智能体框架日益将过程性知识打包为技能:智能体按需读取的指令文件,而公共库如今已包含数千个此类技能。因此,选择读取哪个技能成为策略在 episode 中途做出的决策,但现有信号无法对其进行训练。我们表明,默认的补救措施——对候选集合进行基于结果奖励的强化学习(RL)——无法教会策略做出该决策,原因在于我们识别并命名的结构性问题:选择器信用匮乏。在广播、序列级优势的设定下,命名所选技能的少数 token 在损失中所占份额逐渐消失,且随着轨迹变长,其继承的信用符号愈发错误。即使选择本身是轨迹中最有价值的决策之一,若选择后执行失败,正确的选择也会受到惩罚。对已完成运行的自身训练人工制品的审计证实了这三个特性,每个特性都随视距单调恶化。SkillGate 通过设计消除了该问题:它将 token 支持划分为两个不相交的信用通道,结果信用仅传递给执行 token,而单独的动作局部优势仅传递给恰好命名技能的 token,且仅当轨迹的单次读取是正确的时才为正。在 16 个候选集合下的五个智能体基准测试中,SkillGate 将 9B 策略的试验成功率从 40.8%提升至 53.2%,远超仅使用结果奖励的相同预算,同时将接触误导性候选的概率降低了三分之二,并减少了技能读取次数。

英文摘要

Agent frameworks increasingly package procedural knowledge as skills: instruction files an agent reads on demand, while public libraries now hold thousands of them. Which skill to read has thus become a decision the policy itself makes in the middle of an episode, yet no existing signal trains it. We show that the default remedy, outcome-rewarded RL over the candidate slate, cannot teach it, for a structural reason we identify and name selector credit starvation: under a broadcast, sequence-level advantage, the few tokens that name the chosen skill carry a vanishing share of the loss, and the credit they inherit is increasingly wrong-signed as trajectories lengthen. A correct choice is punished whenever the execution after it fails, even though the choice itself is among the most valuable decisions in the trajectory. Auditing a completed run's own training artifacts confirms all three properties, each worsening monotonically with horizon. SkillGate removes the failure by construction: it partitions the token support into two disjoint credit channels, outcome credit reaching only execution tokens, and a separate action-local advantage reaching exactly the skill-naming tokens, positive only when a trajectory's single read is the correct one. On five agentic benchmarks under a 16-candidate slate, SkillGate lifts a 9B policy from 40.8% to 53.2% trial success, well ahead of the identical budget spent on outcome reward alone, while cutting exposure to misleading candidates by two thirds and reading fewer skills.

Comments17 pages, 6 figures, 6 tables. Code: https://github.com/DeepExperience/SkillGate. Models: https://huggingface.co/simonlqy/SkillGate-9B

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑