arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可达性感知扩散策略优化

Reachability-Aware Diffusion Policy Optimization

Hikmet Simsir, Kutay Demiray, Ozgur S. Oguz

arXiv 2610.05969首次发表:更新:

发表机构

Bilkent University(比尔肯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出无模型的可达性感知扩散策略优化方法,结合预测性首次命中安全估计与累积成本预算反馈,学习折扣首次命中可达性值塑造奖励,通过加权去噪回归改进扩散策略,在十个连续控制安全任务中实现有竞争力的奖励-成本权衡,显著减少约束违反。

AI 中文摘要

扩散策略为连续控制强化学习提供了富有表现力的动作分布。然而,安全感知的在线扩散策略优化仍未得到充分探索,尤其是那些在没有显式动力学模型的情况下利用预测性可达性信息的方法。我们提出了可达性感知扩散策略优化(RADPO),这是一种无模型方法,结合了预测性首次命中安全估计与累积成本预算反馈。RADPO学习一个折扣的首次命中可达性值,该值捕获成本事件的折扣风险,对更早发生的事件赋予更大的权重,并使用该信号来塑造奖励。一个独立的对偶式乘数根据相对于规定预算的实现情节成本来调整塑造强度。扩散行动者通过基于奖励评论家评分的候选动作的加权去噪回归来改进。我们的方法既不需要学习的动力学模型,也不需要通过评论家的动作梯度,更不需要对反向扩散采样器进行微分。我们建立了可达性值的理论性质,并表明累积可达性惩罚提供了未来折扣累积成本的保守替代。在十个连续控制安全任务中,RADPO实现了有竞争力的奖励-成本权衡,在多个任务上相对于比较基线显著减少了约束违反。我们的理论和实证分析支持将可达性与累积预算反馈相结合是安全感知扩散策略的可行方法。

英文摘要

Diffusion policies provide expressive action distributions for continuous-control reinforcement learning. However, safety-aware online diffusion policy optimization remains underexplored, particularly methods that use predictive reachability information without an explicit dynamics model. We propose Reachability-Aware Diffusion Policy Optimization (RADPO), a model-free method that combines predictive first-hit safety estimation with cumulative-cost budget feedback. RADPO learns a discounted first-hit reachability value that captures the discounted risk of a cost event, assigns larger weight to events that occur sooner, and uses this signal to shape the reward. A separate dual-like multiplier adjusts the shaping strength according to realized episodic costs relative to a prescribed budget. The diffusion actor improves through weighted denoising regression on candidate actions scored by the reward critic. Our approach requires neither a learned dynamics model, action gradients through the critics, nor differentiation through the reverse diffusion sampler. We establish theoretical properties of the reachability value and show that accumulated reachability penalty provides a conservative surrogate for future discounted cumulative cost. Across ten continuous-control safety tasks, RADPO achieves competitive reward-cost trade-offs, with substantial reductions in constraint violations on several tasks relative to the compared baselines. Our theoretical and empirical analysis supports that combining reachability with cumulative budget feedback is a viable approach to safety-aware diffusion policies.

CommentsUnder review at ICLR 2027

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑