AI 中文总结
本文提出Fast-TD-MPC,一个自适应路由框架,在快速策略执行与测试时规划间权衡,在103个连续控制任务中保持竞争力,推理加速约4倍,并在扰动下保持鲁棒性。
AI 中文摘要
数据驱动模型预测控制(MPC)将学习到的世界模型与在线轨迹优化相结合,在连续控制中取得了强劲性能。然而,每步采样和评估数百条候选轨迹的成本限制了部署,使其控制频率远低于实时机器人技术的要求。受人类认知双过程理论的启发,该理论区分了快速、直觉性的处理(系统1)与较慢、深思熟虑的推理(系统2),我们提出疑问:是否每个决策都需要相同程度的计算深思?我们提出了Fast-TD-MPC,一个轻量级框架,它在快速策略执行与测试时规划之间自适应地路由,将昂贵的深思保留给最需要的状态。Fast-TD-MPC在103个连续控制任务中提供了具有竞争力的任务性能,同时实现了高达约4倍的推理加速。在外部扰动下,Fast-TD-MPC选择性地回退到规划,保持了与原始规划器相当的鲁棒性。
英文摘要
Data-driven model predictive control (MPC) combines learned world models with online trajectory optimization, achieving strong performance in continuous control. However, the per-step cost of sampling and evaluating hundreds of candidate trajectories restricts deployment to control frequencies well below what real-time robotics demands. Motivated by the dual-process theory of human cognition, which distinguishes between fast, intuitive processing (System 1) and slower, deliberative reasoning (System 2), we ask whether every decision requires the same degree of computational deliberation. We propose Fast-TD-MPC, a lightweight framework that adaptively routes between fast policy execution and test-time planning, reserving costly deliberation for states where it is most needed. Fast-TD-MPC delivers competitive task performance across 103 continuous control tasks while achieving up to ~4x faster inference. Under external disturbances, Fast-TD-MPC selectively falls back to planning, maintaining robustness comparable to the original planner.