发表机构
Yuanbao Team, Tencent; Tsinghua University; Huazhong University of Science and Technology; School of Artificial Intelligence, Shanghai Jiao Tong University(腾讯元宝团队; 清华大学; 华中科技大学; 上海交通大学人工智能学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出图条件化的在线策略蒸馏(GC-OPD),利用图索引教师执行历史丰富评分上下文,在多个基准上显著提升智能体成功率,且无需任务特定优化。
AI 中文摘要
在线策略蒸馏(OPD)通过教师反馈在智能体生成的轨迹上训练紧凑语言智能体。在多轮任务中,复合误差可能使智能体偏离教师有效监督的范围。我们提出图条件化的在线策略智能体蒸馏(GC-OPD),该方法通过执行证据丰富现成教师模型的评分上下文。一个图结构按共享状态索引重复的教师执行过程,同时保留完整的成功与失败历史。在每个智能体回合后,GC-OPD 检索当前状态的参考或历史替代方案,并将其与智能体的后见之明相结合,以对原始思考-行动令牌进行评分。使用相同的原始教师模型,GC-OPD 在 ScienceWorld(4B 智能体)上将平均成功率从 24.70% 提升至 48.78%,在 ALFWorld Unseen 上从 53.36% 提升至 85.26%,在 WebShop 上从 29.10% 提升至 37.65%。在匹配智能体规模的情况下,与使用 GRPO 训练的教师模型的所有评估 OPD 基线相比,GC-OPD 在 ScienceWorld 和 ALFWorld 上也取得了更高的平均成功率;其中最强的 ScienceWorld 4B 基线达到了 46.66%。GC-OPD 不需要针对特定任务进行教师优化。
英文摘要
On-policy distillation (OPD) trains compact language agents with teacher feedback on student-generated trajectories. In multi-turn tasks, compounding errors can move students beyond the teacher's effective supervision. We introduce Graph-Conditioned On-Policy Agent Distillation (GC-OPD), which enriches an off-the-shelf teacher's scoring context with execution evidence. A graph indexes repeated teacher executions by shared states while preserving complete successful and failed histories. After each student episode, GC-OPD retrieves current-state references or historical alternatives and combines them with student hindsight to score the original thought-action tokens. Using the same original teachers, GC-OPD improves mean success over vanilla OPD from 24.70% to 48.78% on ScienceWorld (4B student), from 53.36% to 85.26% on ALFWorld Unseen, and from 29.10% to 37.65% on WebShop. At matched student sizes, it also achieves higher mean success than every evaluated OPD baseline using GRPO-trained teachers on ScienceWorld and ALFWorld; the strongest such ScienceWorld 4B baseline reaches 46.66%. GC-OPD requires no task-specific teacher optimization.
Comments18 pages, 3 figures