发表机构
Wuhan University(武汉大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出SFAC方法,通过条件扩散和Doob h-变换实现KL正则化策略改进,结合极小极大评论家与SNIS估计漂移修正,并给出有限样本分析,在连续控制任务中优于离线初始化。
AI 中文摘要
扩散策略能够表示多模态的动作分布,但优势加权更新并未指明如何从由此产生的目标分布中进行采样。我们提出了Schrödinger--Föllmer Actor--Critic(SFAC),一种通过条件扩散实现Kullback--Leibler(KL)正则化策略改进的离线到在线强化学习(RL)方法。一个极小极大Bellman评论家估计优势函数,该函数定义了一个指数倾斜的目标策略。Doob $h$-变换将此更新表示为对参考扩散漂移的修正。我们推导了该修正的后验均值表示,并使用成对自归一化重要性采样(SNIS)对其进行估计。对这些漂移目标进行监督回归即可更新神经演员,无需评论家的动作梯度。在小更新范围内,KL正则化更新遵循自然策略梯度方向,而Doob修正表示在扩散漂移空间中相同的局部变化。在适当条件下,我们推导出有限样本界,将评论家估计、神经漂移回归、有限样本SNIS、扩散离散化以及继承的演员误差对期望平均策略次优性的影响分离开来。合成实验评估了所提方法对指定优势倾斜目标的逼近精度以及对采样预算的敏感性。在六个离线到在线连续控制任务中,参考锚定实现所获得的最终窗口回报高于其相应离线初始化的回报。
英文摘要
Diffusion policies represent multimodal action distributions, but an advantage-weighted update does not specify how to sample from the resulting target distribution. We propose Schrödinger--Föllmer Actor--Critic (SFAC), an offline-to-online reinforcement learning (RL) method for Kullback--Leibler (KL)-regularized policy improvement through conditional diffusion. A minimax Bellman critic estimates the advantage function, which defines an exponentially tilted target policy. A Doob $h$-transform expresses this update as a correction to the reference diffusion drift. We derive a posterior-mean representation of the correction and estimate it using paired self-normalized importance sampling (SNIS). Supervised regression on these drift targets updates the neural actor without critic action gradients. In the small-update regime, the KL-regularized update follows the natural policy-gradient direction, and the Doob correction represents the same local change in the space of diffusion drifts. Under suitable conditions, we derive finite-sample bounds that separate the effects of critic estimation, neural drift regression, finite-sample SNIS, diffusion discretization, and inherited actor error on expected average policy suboptimality. Synthetic experiments assess the accuracy of approximation to prescribed advantage-tilted targets and sensitivity to sampling budgets. On six offline-to-online continuous-control tasks, a reference-anchored implementation achieves higher final-window returns than those of its corresponding offline initialization.