发表机构
University of British Columbia; Inverted AI; Alberta Machine Intelligence Institute(英属哥伦比亚大学; Inverted AI公司; 阿尔伯塔机器智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究受限生成模型,提出将在线展开获得约束指导纳入训练过程的微调框架,通过微分固定噪声时间表使训练与采样对齐,实验表明该方法能提高约束满足度且保持采样质量。
AI 中文摘要
受限生成模型旨在生成满足复杂可行性约束且忠实于数据分布的样本。现有受限生成方法通常通过训练时优化或采样时校正来强制执行约束。训练时优化方法在训练分布诱导的状态上进行优化,这可能与采样时遇到的状态有很大不同。采样时校正方法则在推理时修改采样过程,引入分布偏移并需要昂贵的调优,特别是对于少步采样。我们提出了一个微调框架,将通过在线展开获得的约束指导纳入训练过程,通过对用于数值积分去噪过程的固定噪声时间表进行微分,使训练与采样对齐。这使模型暴露于去噪轨迹上出现的违反情况,并使扩散学习与采样过程对齐。跨多个任务的实验表明,我们的方法在保持与现有方法具有竞争力的采样质量的同时,提高了约束满足度。
英文摘要
Constrained generative models aim to produce samples that satisfy complex feasibility constraints while remaining faithful to the data distribution. Existing constrained generation methods typically enforce constraints either through training-time optimization or sampling-time correction. Training-time optimization approaches optimize on states induced by the training distribution, which can differ substantially from those encountered during sampling. Sampling-time correction methods instead modify the sampling process at inference, introducing distribution shift and requiring expensive tuning, particularly for few-step sampling. We propose a fine-tuning framework that incorporates constraint guidance obtained through online rollout into the training process, which aligns training with sampling by differentiating through the fixed noise schedule used to numerically integrate the denoising process. This exposes the model to violations that arise along the denoising trajectory and aligns diffusion learning with the sampling process. Experiments across multiple tasks show that our method improves constraint satisfaction while maintaining competitive sampling quality compared to prior methods.