发表机构
The University of Sydney(悉尼大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
LatticeSMC提出预算匹配的分块引导采样器,在块索引和去噪步骤二维格上精确重采样,提升节拍对齐与提示遵循度,并保持生成质量。
AI 中文摘要
面向音乐、动作和视频的长序列生成器逐块生成序列,每个块通过迭代去噪生成,而奖励则在整个序列上定义。现有的推理时引导方法通常一次只作用于一个轴:在末尾进行最佳N采样、在去噪步骤间进行Feynman-Kac引导、或在块间进行流式剪枝,且这些方法常常在计算量不匹配或回报规则不同的情况下进行比较。我们引入了预算匹配的分块引导,并提出了LatticeSMC,这是一种基于块索引和去噪步骤二维格上的Feynman-Kac模型推导出的采样器。两个伸缩结果使其设计精确:对于块加性奖励,两个轴产生相同的权重,因此重采样应发生在前瞻成本最低的地方;对于终端奖励,任何前缀分数都定义了一个精确的中间势,使得前缀可评估奖励的扭曲无需估计或额外的去噪调用。LatticeSMC在块边界处对这些势进行重采样,并且在评分免费时在块内进行重采样,返回加权样本或最佳粒子。在匹配计算量下,在音乐到舞蹈扩散和40秒文本到音乐生成任务中,它将节拍对齐从0.234提升到0.441(最佳N采样为0.354),在32个粒子时将提示遵循度从0.470提升到0.560,同时保持留出集的质量。它还在长程奖励上保持优势,并在60-77%的成对比较中被人类评分者偏好。最后,我们表明承诺强度应遵循当前势中的信息,而前瞻的价值则由未来奖励的集合内可预测性来预测。
英文摘要
Long-form generators for music, motion and video produce sequences chunk by chunk, with each chunk generated by iterative denoising while rewards are defined over the full sequence. Existing inference-time steering methods typically act on one axis at a time: best-of-N at the end, Feynman-Kac steering across denoising steps, or streaming pruning across chunks, and are often compared under unmatched compute or different return rules. We introduce budget-matched chunked steering and propose LatticeSMC, a sampler derived from a Feynman-Kac model on the two-dimensional lattice of chunk index and denoising step. Two telescoping results make its design exact: for chunk-additive rewards, the two axes induce identical weights, so resampling should occur where lookahead is cheapest; for terminal rewards, any prefix score defines an exact intermediate potential, making prefix-evaluable rewards twists with no estimation or extra denoiser calls. LatticeSMC resamples on these potentials at chunk boundaries and, when scoring is free, within chunks, returning either a weighted draw or the best particle. Under matched compute, on music-to-dance diffusion and 40-second text-to-music generation, it raises beat alignment from 0.234 to 0.441 (best-of-N: 0.354) and prompt adherence from 0.470 to 0.560 at 32 particles, while preserving held-out quality. It also retains its advantage on long-range rewards and is preferred by human raters in 60-77 percent of pairwise comparisons. Finally, we show that commitment strength should follow the information in the current potential, while the value of lookahead is predicted by the within-set predictability of future reward.
Comments29 pages, 7 figures