利用推理时计算:通过去噪轨迹的全局调度优化扩散模型
Leveraging Inference-Time Compute for Diffusion Models via Global Scheduling of Denoising Trajectories
查看机构详情
- Peking University(北京大学)
- University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本文提出一种基于全局调度的计算预算分配方法,通过分解预期收益并求解注水结构整数规划,在扩散模型去噪轨迹中优化搜索资源,实验显示可减少20%-50%的函数评估。
中文摘要 AI 辅助
扩散模型通过遍历一条去噪轨迹来生成样本,该轨迹是一系列随机的降噪步骤,将纯噪声转化为目标分布的样本。在部署时,额外的计算可以在不重新训练的情况下提高样本质量:在每一步中,采样器抽取多个候选噪声样本,用称为验证器的质量准则对所得预测进行评分,并保留最佳候选,代价是每个候选进行一次网络评估。这引发了一个资源分配问题:在固定的函数评估预算下,搜索工作应如何分配在去噪轨迹的各个步骤上?我们将此表述为一个计算预算分配问题。首先,我们证明,在步长的一阶近似下,在某一步评估K个候选的预期收益可分解为一个内生的、步骤特定的灵敏度参数乘以一个通用的样本量因子,该因子等于K个标准正态抽取的预期最大值。其次,对于固定的灵敏度分布,最优分配求解一个具有注水结构的可分离凹整数规划;在总灵敏度固定时,其相对于均匀分配的优势随灵敏度在优超序中的离散度增加而增大。第三,我们证明当灵敏度在不同实例间变化时,任何自适应策略都无法避免最坏情况下的遗憾,该遗憾随轨迹长度线性增长,这促使一种设计:离线锚定分配,在线仅调整以恢复实例特定的松弛。我们将分析从独立随机搜索扩展到更广泛的局部搜索算子族,并将其实现为可执行的算法。在三个扩散采样器族上的实验表明,所提出的分配以20%至50%更少的函数评估达到了均匀基准的质量。
英文摘要
Diffusion models generate a sample by traversing a denoising trajectory, a sequence of stochastic noise-reduction steps that transforms pure noise into a draw from a target distribution. At deployment time, additional computation can improve sample quality without retraining: at each step, the sampler draws several candidate noise samples, scores the resulting predictions with a quality criterion called the verifier, and retains the best candidate at the cost of one network evaluation per candidate. This raises a resource allocation question: given a fixed budget of function evaluations, how should search effort be distributed across the steps of the denoising trajectory? We formulate this as a computational budget allocation problem. First, we show that, to leading order in the step size, the expected gain from evaluating $K$ candidates at a step factorizes into an endogenous, step-specific sensitivity parameter times a universal sample-size factor equal to the expected best of $K$ standard-normal draws. Second, for a fixed sensitivity profile, the optimal allocation solves a separable concave integer program with water-filling structure; at fixed total sensitivity, its advantage over uniform allocation increases with sensitivity dispersion in the majorization order. Third, we prove that when sensitivities vary across instances, no adaptive policy can avoid worst-case regret that grows linearly in the trajectory length, which motivates a design that anchors the allocation offline and adapts online only to recover instance-specific slack. We extend the analysis from independent random search to a broader family of local search operators, and instantiate it as an implementable algorithm. Experiments on three families of diffusion samplers show that the proposed allocation attains the quality of the uniform benchmark with 20 to 50 percent fewer function evaluations.