SCAD:面向长时程智能体的结构化信用分配与蒸馏
SCAD: Structured Credit Assignment and Distillation for Long-Horizon Agents
浏览论文内容
中文总结 AI 辅助
SCAD通过结构化信用分配与局部蒸馏,结合终端奖励和教师指导,提升长时程智能体的规划与执行,在文本和多模态任务上分别提升4.48和4.19个百分点。
中文摘要 AI 辅助
训练长时程智能体解决复杂任务需要在整个交互序列上进行有效监督。然而,稀疏的终端奖励掩盖了中间步骤的贡献,而在线策略蒸馏在智能体生成的历史轨迹变长时可能会丢失有价值的教师指导。为解决这一问题,我们提出SCAD,该方法将交互组织为规划与有界子任务执行,在局部上下文中蒸馏执行过程,并通过跨轨迹子任务前缀树细化规划信用,其中规划获得完整终端信用,执行获得正终端信用和教师指导。在所有评估基准上,SCAD在最强训练基线上将文本任务的宏平均准确率提升了4.48个百分点,多模态任务提升了4.19个百分点。SCAD有效结合了基于结果的信用分配与教师引导的蒸馏,以改进长时程智能体的规划与执行。
英文摘要
Training long-horizon agents to solve complex tasks requires effective supervision over extended interaction sequences. However, sparse terminal rewards obscure intermediate contributions, while on-policy distillation can lose informative teacher guidance as student-generated histories grow. To address this problem, we introduce SCAD, which organizes interactions into planning and bounded subtask execution, distills execution in local contexts, and refines planning credit through cross-rollout subtask prefix trees, with planning receiving full terminal credit and execution receiving positive terminal credit and teacher guidance. Across all evaluated benchmarks, SCAD improves macro-average accuracy over the strongest training baseline by 4.48 percentage points for text tasks and 4.19 points for multimodal tasks. SCAD effectively combines outcome-based credit assignment with teacher-guided distillation to improve planning and execution in long-horizon agents.
发表机构
- Beijing University of Posts and Telecommunications(北京邮电大学)
- Tsinghua University(清华大学)
- Nanyang Technological University, Singapore(新加坡南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。