arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过熵反馈实现确定性MPPI的信息论自适应冷却

Information-Theoretic Adaptive Cooling for Deterministic MPPI via Entropy Feedback

Shuqi Wang, Wenrong Sun, Tao Han, Yue Gao, Xiang Yin

arXiv 2607.14245首次发表:更新:

发表机构

School of Automation & Intelligent Sensing, Shanghai Jiao Tong University; Department of Physics, The Hong Kong University of Science and Technology(自动化与智能感知学院,上海交通大学; 物理系,香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究基于MPPI的确定性最优控制中冷却策略难题,提出ITAC框架,利用重要性权重香农熵作反馈信号调节温度,证明其渐近收敛,实验显示该框架提升采样效率且收敛更快,不影响MPPI无导数特性。

AI 中文摘要

本文研究使用模型预测路径积分(MPPI)控制的确定性最优控制,这是一种基于采样且无导数的框架,适用于具有复杂动力学和非光滑目标的系统。在确定性MPPI中,温度必须降至零才能恢复真正的最优解,但设计有效的冷却策略仍是一个基本挑战。现有方法通常依赖预定义的开环策略,限制了算法的效率和鲁棒性。为克服此限制,我们提出信息论自适应冷却(ITAC)框架,它使用重要性权重的香农熵作为在线反馈信号来调节温度。该机制使冷却速率适应当前采样状态,权重分散时快速推进,集中时谨慎冷却。我们证明了所得方案渐近收敛于确定性最优解,并进一步推导了一个临界熵阈值,该阈值可形成防止权重过早崩溃的平滑屏障。在非光滑信号时态逻辑运动规划任务上的实验表明,ITAC提高了采样效率,且在不牺牲MPPI无导数特性的情况下,比现有基线实现了更快的收敛。

英文摘要

This paper investigates deterministic optimal control using Model Predictive Path Integral (MPPI) control, a sampling-based and derivative-free framework well suited for systems with complex dynamics and nonsmooth objectives. In deterministic MPPI, the temperature must be driven to zero to recover the true optimum, yet the design of an effective cooling schedule remains a fundamental challenge. Existing methods typically rely on predefined open-loop schedules, which limit the efficiency and robustness of the algorithm. To overcome this limitation, we propose an Information-Theoretic Adaptive Cooling (ITAC) framework that uses the Shannon entropy of the importance weights as an online feedback signal to regulate the temperature. The proposed mechanism adapts the cooling rate to the current sampling state, enabling fast progress when the weights are diffuse and cautious cooling when they become concentrated. We prove asymptotic convergence of the resulting scheme to the deterministic optimum, and further derive a critical entropy threshold that leads to a smooth barrier against premature weight collapse. Experiments on nonsmooth signal temporal logic motion-planning tasks show that ITAC improves sampling efficiency and achieves substantially faster convergence than state-of-the-art baselines without sacrificing the derivative-free nature of MPPI.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑