发表机构
The University of Tokyo(东京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对离线强化学习中轨迹规划器的问题,提出捷径轨迹规划(STP)框架,将捷径模型作为轨迹生成器,单阶段训练条件捷径轨迹模型,支持可调推理,用增强可行性感知校正的评论家选候选计划,在多任务基准测试中性能强且简化训练管道。
AI 中文摘要
基于扩散的轨迹规划器在离线强化学习中表现出色,但其迭代去噪过程推理成本高。基于一致性的规划器减少了采样步骤,但依赖两阶段师生蒸馏管道,增加训练成本并可能引入不稳定性。我们提出捷径轨迹规划(STP),一个基于离线模型的强化学习框架,将捷径模型作为高效轨迹生成器。STP在单阶段训练条件捷径轨迹模型,通过步长条件支持可调的单步和多步推理,并使用增强了可行性感知校正的评论家选择候选计划。在包括运动、导航、操纵和灵巧控制任务的标准D4RL基准测试中,STP在简化快速生成规划训练管道的同时取得了强大性能。
英文摘要
Diffusion-based trajectory planners have shown strong performance in offline reinforcement learning, but their iterative denoising process often incurs high inference cost. Consistency-based planners reduce the number of sampling steps, yet they typically rely on a two-stage teacher--student distillation pipeline that increases training cost and may introduce instability. We propose Shortcut Trajectory Planning (STP), an offline model-based reinforcement learning framework that incorporates shortcut models as efficient trajectory generators. STP trains a conditional shortcut trajectory model in a single stage, supports adjustable one-step and few-step inference through step-size conditioning, and selects candidate plans using a critic augmented with feasibility-aware correction. Across standard D4RL benchmarks, including locomotion, navigation, manipulation, and dexterous control tasks, STP achieves strong performance while simplifying the training pipeline for fast generative planning.
Comments16 pages, 3 figures