AI 中文总结
该研究针对扩散策略连续控制的计算瓶颈,提出POGP框架,通过学习前缀值函数实现自适应去噪停止,减少约2.7倍迭代次数,同时提升任务性能。
AI 中文摘要
扩散策略是连续控制中强大的策略类别,但其迭代去噪过程会产生显著的计算瓶颈。降低该成本需要根据每个动作的难度调整去噪步数,同时保持任务性能。我们提出前缀最优生成策略(Prefix-Optimal Generative Policies,POGP),该框架通过在去噪链上应用贝尔曼风格的递归,在每个中间去噪步骤学习前缀值函数。前缀值函数有两个用途:一是提供辅助训练目标,鼓励中间输出成为高质量动作;二是支持测试时的停止规则,当额外步骤不太可能产生有意义的改进时终止去噪。在四个MuJoCo环境中与12个基准方法对比,POGP将所需的去噪迭代次数减少了约2.7倍,同时保持接近完整的任务性能。与最先进的动态扩散基准方法相比,前缀训练还将最终任务性能提高了约3.5%。这些结果表明,监督中间去噪步骤不仅对自适应提前停止有用,而且作为辅助目标也能改进学习到的策略。
英文摘要
Diffusion policies are a powerful policy class for continuous control, but their iterative denoising process creates a substantial computational bottleneck. Reducing this cost requires adapting the number of denoising steps to the difficulty of each action while preserving task performance. We introduce Prefix-Optimal Generative Policies (POGP), a framework that learns a prefix value function at every intermediate denoising step through a Bellman-style recursion over the denoising chain. The prefix value function serves two purposes: it provides an auxiliary training objective that encourages intermediate outputs to become high-quality actions, and it enables a test-time stopping rule that terminates denoising when additional steps are unlikely to produce meaningful improvement. Across four MuJoCo environments and comparisons with 12 baselines, POGP reduces the required number of denoising iterations by approximately 2.7-fold while retaining near-full task performance. Compared with state-of-the-art dynamic diffusion baselines, prefix training also improves final task performance by approximately 3.5%. These results indicate that supervising intermediate denoising steps is useful not only for adaptive early stopping, but also as an auxiliary objective that improves the learned policy.