arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ART 用于扩散采样:连续时间控制与 Actor-Critic 学习

Adaptive Reparametrized Time for Score-Based Diffusion Sampling

Yilie Huang, Wenpin Tang, Xun Yu Zhou

arXiv 2607.02137首次发表:更新:

发表机构

Department of Applied Mathematics, The Hong Kong Polytechnic University; Department of Industrial Engineering and Operations Research, Columbia University; Data Science Institute, Columbia University(香港理工大学应用数学系; 哥伦比亚大学工业工程与运筹学系; 哥伦比亚大学数据科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出自适应重参数化时间 (ART) 方法,将扩散采样的时间步分配建模为连续时间控制问题,并通过强化学习求解,以自适应学习采样时间表,提升样本质量。

AI 中文摘要

我们研究了基于分数的扩散采样的时间步分配问题,其中学习到的逆向时间动力学在有限网格上离散化。均匀和手工设计的调度是标准选择,但它们依赖于固定预设,因此可能次优。为解决这一限制,我们提出自适应重参数化时间 (ART),一种连续时间控制公式,通过将采样时钟的速度视为控制来学习时间变化,使得学习到的时钟上的均匀网格在原始扩散时间中诱导出自适应时间步。基于前导阶欧拉误差代理,ART 为沿采样轨迹分配时间步提供了一个有原则的目标。为解决这一确定性控制问题,我们引入 ART-RL,一种带有高斯策略的辅助随机化公式,将调度学习转化为连续时间强化学习问题。我们证明了随机化的 ART-RL 公式在优化器层面与 ART 等价,即其最优高斯策略通过其均值恢复最优 ART 时间扭曲率。我们进一步建立了策略评估和策略改进的特征,并推导了基于轨迹的矩恒等式,从而得到可实现的 actor-critic 更新来学习调度。在从受控低维设置到图像生成的实验中,ART-RL 可以通过仅改变时间步网格插入现有扩散采样器,在匹配预算下持续改善样本质量,同时保持采样管道的其余部分不变。学习到的调度还表现出广泛的泛化能力,无需重新训练即可跨采样预算、数据集、求解器、管道和表示空间迁移。

英文摘要

We study timestep allocation for score-based diffusion sampling, where a learned reverse-time dynamics is discretized on a finite grid. Uniform and hand-crafted schedules are standard choices, but they rely on fixed prescriptions and can therefore be suboptimal. To address this limitation, we propose Adaptive Reparameterized Time (ART), a continuous-time control formulation that learns a time change by treating the speed of the sampling clock as the control, so that a uniform grid on the learned clock induces adaptive timesteps in the original diffusion time. Based on a leading-order Euler error surrogate, ART provides a principled objective for allocating timesteps along the sampling trajectory. To solve this deterministic control problem, we introduce ART-RL, an auxiliary randomized formulation with Gaussian policies that turns schedule learning into a continuous-time reinforcement learning problem. We prove that the randomized ART-RL formulation is equivalent to ART at the optimizer level, in the sense that its optimal Gaussian policy recovers the optimal ART time-warping rate through its mean. We further establish policy evaluation and policy improvement characterizations and derive trajectory-based moment identities that yield implementable actor--critic updates for learning the schedule. Across experiments ranging from controlled low-dimensional settings to image generation, ART-RL can be plugged into existing diffusion samplers by changing only the timestep grid, consistently improving sample quality over strong baseline schedules at matched budgets while leaving the rest of the sampling pipeline unchanged. The learned schedules also exhibit broad generalization, transferring without retraining across sampling budgets, datasets, solvers, pipelines, and representation spaces.

Comments37 pages, 14 figures, 8 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑