发表机构
Applied AI Institute(应用人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出TSBM后训练方法,通过交替优化和伴随匹配,将预训练薛定谔桥微调至奖励倾斜目标,在MNIST和CelebA上验证了无配对翻译性能。
AI 中文摘要
薛定谔桥为无配对域翻译提供了熵正则化框架和原则性解决方案。在实践中,预训练的桥可能需要通过奖励来适应人类偏好或物理约束,这一问题与扩散模型中的奖励倾斜密切相关,但在薛定谔桥中尚未得到充分探索。我们引入了倾斜薛定谔桥匹配(TSBM),这是一种后训练方法,用于将源分布$p_0$和目标分布$p_1$之间学习到的桥$P$微调至奖励倾斜的目标$p_1^r\propto p_1e^r$,同时保持源分布$p_0$不变。我们将这种适应过程表述为从$P$初始化的交替优化,提供了理论依据,并推导出基于伴随匹配的实用算法。我们在MNIST上的数字属性和CelebA上的面部属性的无配对图像到图像翻译任务上评估了TSBM。
英文摘要
Schrödinger bridges provide an entropy-regularized framework and a principled solution for unpaired domain translation. In practice, a pretrained bridge may need to be adapted to human preferences or physical constraints through a reward a problem closely related to reward tilting in diffusion models but underexplored for Schrödinger bridges. We introduce Tilted Schrödinger Bridge Matching (TSBM), a post-training method for fine-tuning a learned bridge $P$ between source $p_0$ and target $p_1$ toward a reward-tilted target $p_1^r\propto p_1e^r$, while preserving source $p_0$. We formulate this adaptation as alternating optimization initialized from $P$, provide theoretical justification, and derive a practical algorithm based on Adjoint Matching. We evaluate TSBM on unpaired image-to-image translation targeting digit properties in MNIST and facial attributes in CelebA.