学习根据任务进度行动:从紧凑的教师监督中蒸馏小型智能体
Learning to Act with Task Progress: Distilling Small Agents from Compact Teacher Supervision
浏览论文内容
中文总结 AI 辅助
提出任务进度蒸馏(TPD),通过紧凑的教师监督训练小型智能体,在ALFWorld上以较少示范实现高成功率,显式阶段信息在中等数据量下提供额外提升。
中文摘要 AI 辅助
从大型模型的示范中学习,提供了一种训练小型智能体的方法,使其能够完成重复性任务,而无需在每一步都调用大型模型。一个核心的设计选择是从包含推理、行动和任务进度信息的教师轨迹中保留什么。我们引入了任务进度蒸馏(TPD),一种离线方法,为每个示范动作配对一个描述当前任务阶段的简短标签。学生模型学习这些紧凑的目标,并通过联合评分可行的阶段-动作对来选择动作,由确定性执行器在环境中执行。在ALFWorld上,使用404个示范训练的1.7B学生模型,在TPD或仅动作监督下,平均未见任务成功率达到72.4%,而使用受限动作选择的推理训练学生模型为48.3%。在200个示范时,显式阶段提供了额外的好处,将成功率从仅动作监督的48.0%提高到67.7%。随着示范数量的增加,仅动作的学生模型缩小了差距,两种方法在808个示范时均达到76.9%。共享历史分析将TPD的部分局部优势归因于在子目标之间移动时更好的决策,特别是从对象获取到处理。这些结果表明,紧凑的监督可以训练有效的小型任务智能体,而显式的任务进度在中等示范预算下提供了额外的指导。
英文摘要
Learning from large-model demonstrations offers a way to train small agents that can complete recurring tasks without calling a large model at every step. A central design choice is what to retain from teacher trajectories that contain reasoning, actions, and information about task progress. We introduce Task-Progress Distillation (TPD), an offline approach that pairs each demonstrated action with a short label describing the current task stage. The student learns these compact targets and selects actions by jointly scoring admissible stage--action pairs, which a deterministic harness executes in the environment. On ALFWorld, a 1.7B student trained with 404 demonstrations achieves 72.4\% mean unseen task success with either TPD or action-only supervision, compared with 48.3\% for a reasoning-trained student using constrained action selection. Explicit stages provide an additional benefit at 200 demonstrations, improving success from 48.0\% to 67.7\% over action-only supervision. With more demonstrations, the action-only student closes the gap, and both approaches reach 76.9\% at 808 demonstrations. Shared-history analyses link part of TPD's local advantage to better decisions when moving between subgoals, particularly from object acquisition to processing. These results show that compact supervision can train effective small task agents, while explicit task progress provides additional guidance at an intermediate demonstration budget.
发表机构
- City University of Hong Kong(香港城市大学)
- Shenzhen Loop Area Institute(深圳河套学院)
机构由 AI 辅助整理,请以论文原文为准。