用于使用工具的智能体的持久教师锚定
Persistent Teacher Anchoring for Tool-Using Agents
浏览论文内容
中文总结 AI 辅助
针对现有工具使用智能体蒸馏方法未覆盖工具执行的问题,提出持久教师锚定(PTA),结合块级验证与轮次级承诺,配合持久前瞻提升效率,在两类任务中显著优于OPKD。
中文摘要 AI 辅助
知识蒸馏是大型语言模型(LLM)后训练中的常用技术,其中在线策略知识蒸馏(OPKD)利用智能体生成的轨迹来让学生模型为下游强化学习(RL)做准备。在每个状态下,学生模型需匹配教师模型提供的下一个token分布。当智能体的探索进入教师模型不会访问的状态时,师生分布的差距会不断累积。在工具使用场景中,这种差距会产生严重影响,因为学生模型编写的工具调用会在获得监督前执行,其观察结果会影响后续的上下文序列。提案-验证生成方法通过让教师模型决定保留学生提案的哪些文本来解决这种分布漂移问题,但现有方法仅针对文本生成,未涵盖工具执行环节。我们提出持久教师锚定(PTA),这是一种由学生模型诱导但由教师模型确定的探索过程构建方法。PTA保留了块级验证机制,并新增了轮次级承诺机制,确保工具调用需经教师模型验证完整个轮次后才会发送至环境。我们将已验证的块视为原子生成单元,引入持久前瞻机制,在固定验证器下,通过提前生成未来样本并在学生模型更新时保留未完成的样本,填补探索过程的空闲容量。在Search-R1风格的检索任务和DeepEyes风格的感知RL任务中,在下游RL前应用PTA,在相同下游RL预算下,相比OPKD分别将macro best@4指标提升了2.5和2.8个百分点;而前瞻机制则使吞吐量提升了24%。
英文摘要
Distillation is common in LLM post-training, where on-policy knowledge distillation (OPKD) uses student-generated trajectories to prepare the student for downstream RL. At each state, the student matches a next-token distribution supplied by the teacher. As the rollout enters states the teacher would not visit, the teacher-student distribution gap can accumulate. In tool use, this gap becomes consequential because student-written calls execute before supervision and their observations shape later prefixes. Proposer-verifier generation addresses this drift by letting the teacher decide which student-proposed text is retained during generation. Existing formulations govern text but leave tool execution outside their scope. We propose Persistent Teacher Anchoring (PTA), a student-induced but teacher-committed rollout construction. PTA retains chunk-level verification and adds turn-level commitment, allowing a call to reach the environment only after the teacher has verified the entire turn. Treating verified chunks as atomic generation units, we introduce persistent lookahead, which fills idle rollout capacity by advancing future samples and carrying unfinished ones across student updates under the fixed verifier. Across Search-R1-style retrieval and DeepEyes-style perception RL, applying PTA before downstream RL improves macro best@4 by 2.5 and 2.8 points over OPKD under the same downstream RL budget, while lookahead improves throughput by 24%.
发表机构
- Sogang University(西江大学)
- University of Michigan(密歇根大学)
机构由 AI 辅助整理,请以论文原文为准。