arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于高效离线强化学习的捷径轨迹规划

Shortcut Trajectory Planning for Efficient Offline Reinforcement Learning

Guanquan Wang, Yoshimasa Tsuruoka

arXiv 2607.09336首次发表:更新:

发表机构

The University of Tokyo(东京大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对离线强化学习中轨迹规划器的问题,提出捷径轨迹规划(STP)框架,将捷径模型作为轨迹生成器,单阶段训练条件捷径轨迹模型,支持可调推理,用增强可行性感知校正的评论家选候选计划,在多任务基准测试中性能强且简化训练管道。

AI 中文摘要

基于扩散的轨迹规划器在离线强化学习中表现出色,但其迭代去噪过程推理成本高。基于一致性的规划器减少了采样步骤,但依赖两阶段师生蒸馏管道,增加训练成本并可能引入不稳定性。我们提出捷径轨迹规划(STP),一个基于离线模型的强化学习框架,将捷径模型作为高效轨迹生成器。STP在单阶段训练条件捷径轨迹模型,通过步长条件支持可调的单步和多步推理,并使用增强了可行性感知校正的评论家选择候选计划。在包括运动、导航、操纵和灵巧控制任务的标准D4RL基准测试中,STP在简化快速生成规划训练管道的同时取得了强大性能。

英文摘要

Diffusion-based trajectory planners have shown strong performance in offline reinforcement learning, but their iterative denoising process often incurs high inference cost. Consistency-based planners reduce the number of sampling steps, yet they typically rely on a two-stage teacher--student distillation pipeline that increases training cost and may introduce instability. We propose Shortcut Trajectory Planning (STP), an offline model-based reinforcement learning framework that incorporates shortcut models as efficient trajectory generators. STP trains a conditional shortcut trajectory model in a single stage, supports adjustable one-step and few-step inference through step-size conditioning, and selects candidate plans using a critic augmented with feasibility-aware correction. Across standard D4RL benchmarks, including locomotion, navigation, manipulation, and dexterous control tasks, STP achieves strong performance while simplifying the training pipeline for fast generative planning.

Comments16 pages, 3 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑