arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26585cs.LGcs.CEq-bio.QMstat.ML

GRAS:用于离散扩散模型无训练奖励对齐的引导式降方差提议与自适应选择

VGAS: Variance-Reduced Guidance and Adaptive Selection for Training-Free Reward Alignment in Discrete Diffusion

Kwanyoung Kim

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出GRAS方法,通过降方差提议与自适应重采样温度改进离散扩散模型无训练奖励对齐,在调控DNA和蛋白质设计任务中表现优于现有无训练方法,效果接近或超越奖励微调模型。

中文摘要 AI 辅助

离散扩散模型已成为序列数据的一类强大且被广泛采用的生成器,在推理阶段无需任何重新训练即可引导其适配下游奖励的需求日益重要。这类无训练引导可通过梯度引导、搜索或两者结合来实现。本文研究了结合引导的场景,发现其常规运行存在两个缺陷:引导式提议从单个噪声样本估计梯度,而搜索随后以固定温度重采样粒子,忽略了奖励在每个去噪步骤中的分布情况。我们通过一组不增加去噪器成本的小幅修改来解决这两个问题:对于提议,针对可微奖励采用Rao-Blackwell化估计器以降低估计方差,针对不可微奖励采用留一法基线;对于搜索,我们将每步值标准化为组相对优势,并证明其可简化为单一有效成分——自适应重采样温度。我们将所得方法命名为引导式降方差提议与自适应选择(GRAS)。GRAS简单却有效:在调控DNA和蛋白质设计任务中,它实现了最优的无训练奖励表现,优于现有无训练方法,且与奖励微调模型的表现相当或更优,即使针对不可微奖励也保持有效性。

英文摘要

Masked discrete diffusion models perform strongly on text, code, and biological sequences, but their training objective rewards only naturalness, and retraining the generator for every new reward is expensive. Inference-time steering of a frozen model either guides the sampler by the reward gradient or searches over several trajectories, and recent samplers combine the two. Such combinations are assembled as pipelines that leave three choices at their defaults: a guidance estimate resting on one Gumbel draw per sample, a reward tilting placed without reference to the distribution the combination then targets, and a selection temperature held fixed although the spread of per-step rewards drifts. We identify that distribution and settle the three choices against it. We therefore propose Variance-reduced Guidance and Adaptive Selection (VGAS), a simple yet effective inference-time framework that reduces the variance of the guidance estimate for both reward types, applies the reward tilting in the clean-token logits, where the pretrained schedule is preserved, and sets the selection temperature per step. Across regulatory DNA, protein and small-molecule benchmarks, VGAS attains the best training-free reward and matches or surpasses a reward-fine-tuned generator.

发表机构

  • Gwangju Institute of Science and Technology (GIST)(光州科学技术院)

机构由 AI 辅助整理,请以论文原文为准。

↑