基于嵌套序贯蒙特卡洛的离散扩散推理时控制
Discrete Diffusion Inference-Time Control with Nested Sequential Monte Carlo
浏览论文内容
中文总结 AI 辅助
该研究针对离散扩散语言模型文本生成的推理时控制问题,提出NSMC与FA-NSMC方法,修正了现有粒子方法的缺陷,在毒性和流畅度引导任务上表现优于n-best采样与bootstrap SMC。
中文摘要 AI 辅助
我们研究离散扩散语言模型中文本生成的推理时控制,目标是在不重新训练的情况下引导采样朝向序列级奖励。该领域现有工作聚焦于基于粒子的方法,如n-best采样和自举序贯蒙特卡洛(bootstrap SMC),但分别存在过度乐观和权重退化问题。我们使用嵌套序贯蒙特卡洛方法解决这些局限,针对费曼-卡茨(Feynman--Kac)引导提出嵌套序贯蒙特卡洛(NSMC)和完全适配嵌套序贯蒙特卡洛(FA-NSMC),识别并修正了现有公式中导致最终估计有偏的错误。我们在毒性和流畅度引导任务上评估这些方法,结果显示NSMC和FA-NSMC始终优于n-best采样和bootstrap SMC。
英文摘要
We study inference-time control for text generation in discrete diffusion language models, where the goal is to steer sampling toward sequence-level rewards without retraining. Prior work in this domain has focused on particle-based methods such as best-of-$n$ sampling and bootstrap sequential Monte Carlo, which may suffer from overoptimism and weight degeneracy, respectively. We address these limitations using \emph{nested} sequential Monte Carlo methods. We formulate nested SMC (NSMC) and fully-adapted nested SMC (FA-NSMC) for Feynman--Kac steering, identifying and correcting errors in prior formulations that lead to biased final estimates. We evaluate these methods on toxicity and fluency steering tasks, showing that NSMC and FA-NSMC consistently outperform best-of-$n$ and bootstrap SMC.
发表机构
- University of Amsterdam(阿姆斯特丹大学)
机构由 AI 辅助整理,请以论文原文为准。