arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

奖励传输:通过噪声空间对齐在流匹配中进行属性控制

Reward Transport: Property Control in Flow Matching via Noise-Space Alignment

Kehan Guo, Yili Shen, Yujun Zhou, Yue Huang, Chujie Gao, Shiyi Du, Xiangliang Zhang

arXiv 2607.08781首次发表:更新:

发表机构

University of Notre Dame; Carnegie Mellon University(圣母大学; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究通过奖励传输在流匹配中实现属性控制,训练时用最优传输耦合将噪声空间坐标与分子奖励对齐,推理时改变坐标引导生成分布,实验表明可对logP和QED进行有效控制,且该接口与其他方法互补,明确了耦合级对齐不存在的情况。

AI 中文摘要

流匹配中的耦合通常被视为一种计算选择。本文表明这种耦合可作为对齐接口,通过根据目标分子属性匹配噪声和数据,将可控结构直接嵌入到学习的流场中。在此基础上引入奖励传输,训练时用最优传输耦合将标量噪声空间坐标与分子奖励对齐,推理时改变该坐标可引导生成分布,无需预言机等。在耦合保持极限下,阈值化该坐标可恢复交叉熵方法的截断奖励分布,提供可连续调节的分布级控制旋钮。实验表明在ZINC - 250K和GuacaMol上,扫描标量可对logP进行单调控制并在其操作范围内实现一致的QED控制,同一旋钮对不同目标产生相反结构响应,排除了通用尺寸偏差。该接口与无分类器引导和条件流匹配互补,epsilon预测扩散下的负面结果明确了耦合级对齐在结构上不存在的地方。

英文摘要

The coupling in flow matching -- the rule pairing noise vectors with data points -- is typically treated as a computational choice. We show that this coupling can instead serve as an alignment interface: by matching noise and data according to a target molecular property, it embeds controllable structure directly into the learned flow field. Building on this view, we introduce Reward Transport, which uses optimal transport coupling at training time to align a scalar noise-space coordinate with molecular rewards; at inference, varying this coordinate steers the generated distribution without requiring an oracle, reward model, gradient guidance, or additional computation. In the coupling-preserving limit, thresholding this coordinate recovers the Cross-Entropy Method's truncated reward distribution, providing a principled, continuously adjustable distribution-level control knob. Empirically, on ZINC-250K and GuacaMol, sweeping the scalar induces monotone control of logP and consistent QED control over its operating range; most tellingly, the same knob produces opposite structural responses for different targets, growing molecules for logP but shrinking them for QED, which rules out a generic size bias. The interface is complementary to classifier-free guidance and conditional flow matching, while a negative result under epsilon-prediction diffusion clarifies where coupling-level alignment is structurally absent. Code: https://github.com/KehanGuo2/reward-transport

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑