hint$^2$:用于推理时态逻辑引导的分层世界模型
Composing Learned Robot Behaviors with Temporal Logic at Runtime
浏览论文内容
中文总结 AI 辅助
本文提出hint$^2$,一种用分层世界模型引导策略满足LTL规范的方法,在CALVIN基准及真实UR5e机械臂实验中,其性能优于现有推理引导方法。
中文摘要 AI 辅助
机器人学习的核心目标是让机器人执行运行时指定的丰富指令。大规模语言条件策略在该目标上取得了显著进展,但仍难以处理时间结构与安全约束。线性时态逻辑(Linear Temporal Logic,LTL)是表达复杂非马尔可夫指令的强大语言,然而引导已学习的操纵策略满足LTL要求仍具挑战性:现代策略生成短周期动作块并在闭环中重规划,而几乎所有LTL规范需在长周期轨迹上评估。本文提出hint$^2$,一种在推理时利用分层世界模型引导短周期策略满足复杂LTL规范的方法。核心思路是利用每个世界模型的抽象级别推导两个独立引导目标:高层模型预测动作诱导的任务相关原子命题的未来转移,以引导通过LTL自动机的进度;低层动力学模型预测即时状态演化,以实现精确的局部安全引导。结果表明,hint$^2$克服了当前LTL引导扩散方法的局限,在CALVIN基准上优于现有推理时引导方法,能比语言条件策略更优雅地完成含复杂活性与安全约束的指令,最终在真实UR5e机械臂上验证了hint$^2$处理复杂指令的能力。
英文摘要
Executing Linear Temporal Logic (LTL) instructions with learned robot policies faces two practical challenges: demonstrations may cover individual behaviors without containing the temporal compositions requested at deployment, and semantic success predicates often provide little useful motor guidance through quantitative robustness. We address these challenges by separating behavior learning from temporal composition. From the same offline demonstrations, we train an unconditioned multimodal diffusion policy and a semantic predictor that estimates which semantic outcomes are likely to follow a proposed action sequence. These predictions provide a learned guidance signal without requiring hand-designed robustness measures for semantic task objectives. At deployment, an LTL automaton tracks instruction progress and scores the policy's proposed actions according to their predicted semantic outcomes. Neither learned component receives the specification during training, allowing demonstrated behaviors to be reused in longer, previously unseen temporal compositions without retraining. For additional runtime safety constraints with informative continuous margins, an optional robustness-based controller locally refines the selected actions. Experiments in a navigation environment, CALVIN manipulation, and on a real robot demonstrate reliable execution of complex temporal instructions and additional safety constraints, with substantial improvements over temporal-logic baselines.
发表机构
- Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。