arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StageGuard:通过智能体蒸馏学习长时程机器人任务的阶段转换

StageGuard: Learning Stage Transitions for Long-Horizon Robot Tasks via Agentic Distillation

Jinbang Huang, Yuanzhao Hu, Zhiyuan Li, Ran Qi, Yixin Xiao, Yangzheng Wu, Tengyue Ba, Zhanguang Zhang, Yingxue Zhang

arXiv 2609.20791首次发表:更新:

发表机构

University of British Columbia; University of Toronto; Labs; McGill University(不列颠哥伦比亚大学; 多伦多大学; 2012实验室; 麦吉尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对长时程机器人任务中阶段转换决策困难的问题,提出StageGuard框架,通过智能体蒸馏将教师模型推理转化为轻量级学生VLM的自我解释,实现高效准确的阶段转换预测,并在基准和真实机器人上验证了显著改进。

AI 中文摘要

分层规划框架将多个机器人控制策略的技能组合起来,用于长时程任务执行,其中确定何时终止当前技能并推进到下一个子任务至关重要。现有方法通常依赖预先设计的完成信号检查器,这些检查器在真实世界执行中难以获得。大规模视觉语言模型(VLM)提供了强大的推理能力,但其决策边界与任务完成标准并不天然一致,而云端部署和冗长的推理过程会引入大量延迟,限制了实时监控。我们提出StageGuard,一个用于准确高效阶段转换决策的智能体蒸馏框架。StageGuard结合教师模型推理与演示轨迹,生成关于子任务完成和策略切换的结构化解释。一个轻量级学生VLM利用这些解释生成紧凑的自我解释,用于监督微调。我们在两个基准的轨迹上评估阶段转换预测,并通过集成到BEHAVIOR-1K上的分层机器人控制来评估闭环任务成功率,同时在真实机器人上进一步验证。结果表明,阶段转换预测有显著改进,同时支持高效的在线监控。

英文摘要

Hierarchical planning frameworks combine skills from multiple robot control policies for long-horizon task execution, where determining when to terminate the current skill and advance to the next subtask is essential. Existing approaches often rely on pre-designed completion signal checkers that are hard to obtain in real-world execution. Large-scale vision-language models (VLMs) offer strong reasoning capabilities, but their decision boundaries are not inherently aligned with task completion criteria, while cloud deployment and lengthy reasoning introduce substantial latency, limiting real-time monitoring. We propose StageGuard, an agentic distillation framework for accurate and efficient stage-transition decisions. StageGuard combines teacher-model reasoning with demonstration trajectories to generate structured explanations of subtask completion and policy switching. A lightweight student VLM uses these explanations to generate compact self-explanations, which are used for supervised fine-tuning. We evaluate stage-transition prediction on trajectories from two benchmarks and assess closed-loop task success through integration into hierarchical robot control on BEHAVIOR-1K, with further validation on real robots. Results show substantial improvements in stage-transition prediction while supporting efficient online monitoring.

Comments8 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑