arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35924cs.AI

Grab a Coffee: 具有编译目标的离散扩散模型的未来感知引导

Grab a Coffee: Future-Aware Guidance for Discrete Diffusion with Compiled Objectives

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • University of California, Los Angeles(加州大学洛杉矶分校)
  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

Hua, Xu, Dongxin Li, Gwen Yidou-Weng, Guy Van den Broeck, Wei Wang, Anji Liu

AI总结:

COFFEE框架通过分离序列依赖性与目标,避免枚举补全,将全局偏好转移到未解析位置,实现离散扩散模型的高效引导,无需重训练,并在多基准上取得强控制效果。

AI中文摘要:

离散扩散模型通过并行迭代解析多个标记来生成序列,为从左到右的生成提供了一种灵活的替代方案。然而,使用序列级目标来引导这一过程是困难的,因为一个未解析标记的值取决于它可以与之形成高奖励序列的其他标记。枚举所有此类补全会使整个引导计算随着未解析位置的数量呈指数增长。我们提出了COFFEE,一种即插即用的框架,通过将序列依赖性与目标分离来避免这种枚举。在每个扩散步骤中,一个无目标载体吸收由去噪器预测的边际标记分布,以构建未解析标记上的联合模型,而一个编译的有限状态模型记录它们的组合如何影响序列级偏好。通过配对它们的状态,COFFEE能够将全局偏好转移到未解析位置,并在不重新训练扩散模型的情况下采样干净的重建。同一框架支持显式硬约束和学习到的软目标。我们在多个符号、语言和生物基准上评估了COFFEE,它在任务相关的质量和多样性权衡下取得了强大的控制结果。通过使目标可用于推理而不仅仅是评估,COFFEE将联合条件、补全加权引导和基于优化的约束引入预训练的神经生成中,展示了神经符号方法在扩散引导中的潜力。

英文摘要:

Discrete diffusion models generate sequences by iteratively resolving multiple tokens in parallel, offering a flexible alternative to left-to-right generation. However, guiding this process with a sequence-level objective is difficult because the value of one unresolved token depends on the other tokens with which it can form a high-reward sequence. Enumerating all such completions makes the whole guidance computation grow exponentially with the number of unresolved positions. We introduce COFFEE, a plug-and-play framework that avoids this enumeration by separating sequence dependence from the objective. At each diffusion step, a target-free carrier absorbs the marginal token distributions predicted by the denoiser to construct a joint model over the unresolved tokens, while a compiled finite-state model records how their combinations affect the sequence-level preference. Pairing their states allows COFFEE to transfer global preferences to unresolved positions and sample a clean reconstruction without retraining the diffusion model. The same framework supports explicit hard constraints and learned soft objectives. We evaluate COFFEE across multiple symbolic, language, and biological benchmarks, where it achieves strong control results with task-dependent quality and diversity trade-offs. By making objectives available to inference rather than only evaluation, COFFEE brings joint conditioning, completion-weighted guidance, and optimization-based constraints into pretrained neural generation, showing the potential of neural-symbolic methods in diffusion guidance.

↑