arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FestDPO:基于直接偏好优化的少步生成器对齐

FestDPO: Few-step Generator Alignment with Direct Preference Optimization

Jaewoo Lee, Kyuil Sim, Hyeongyu Kang, Kanghoon Lee, Woocheol Shin, Jinkyoo Park

arXiv 2609.34673首次发表:更新:

发表机构

KAIST; MongooseAI; Omelet(韩国科学技术院; MongooseAI; Omelet)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

FestDPO通过非参数似然估计将直接偏好优化扩展到少步生成模型,实现样本级对齐,在文本到图像和蛋白质骨架生成任务中优于基线。

AI 中文摘要

少步生成模型能够在少量函数评估内生成高保真样本。尽管具有这种效率,生成的样本可能不具备理想的属性。当这些属性难以编码为显式奖励函数时,直接偏好优化(DPO)可以利用成对偏好反馈来对齐生成模型,而无需训练单独的奖励模型。然而,将DPO扩展到少步生成模型具有挑战性,因为少步生成模型通常是隐式的,使得DPO所需的似然评估难以处理。为了应对这一挑战,我们引入了少步DPO(FestDPO),这是DPO对少步生成模型的扩展,利用来自经验样本的非参数似然估计。通过利用少步生成模型的快速采样能力,我们的方法使基于样本的DPO损失近似在计算上可行。此外,基于样本的公式使FestDPO对模型家族和采样过程具有不可知性。我们的玩具实验表明,FestDPO在四个少步生成器上匹配了奖励倾斜的目标分布。对于现实世界任务,我们在两个领域评估了FestDPO:文本到图像生成和蛋白质骨架生成。在文本到图像生成中,FestDPO在对抗基础模型的胜率和人类评估分数方面均优于偏好优化基线。在蛋白质骨架生成中,它实现了比基线更高的β-折叠比例和更好的结构可设计性。

英文摘要

Few-step generative models can generate high-fidelity samples within a few function evaluations. Despite this efficiency, generated samples may not exhibit desirable properties. When these properties are difficult to encode as an explicit reward function, direct preference optimization (DPO) can align generative models using pairwise preference feedback without training a separate reward model. However, extending DPO to few-step generative models is challenging because few-step generative models are generally implicit, making the likelihood evaluation required by DPO intractable. To address this challenge, we introduce Few-step DPO (FestDPO), an extension of DPO for few-step generative models that leverages nonparametric likelihood estimation from empirical samples. By exploiting the fast sampling capabilities of few-step generative models, our approach makes sample-based approximation of DPO loss computationally feasible. Furthermore, the sample-based formulation makes FestDPO agnostic to the model family and sampling procedure. Our toy experiment demonstrates that FestDPO matches the reward-tilted target distribution across four few-step generators. For real-world tasks, we evaluate FestDPO in two domains: text-to-image generation and protein backbone generation. In text-to-image generation, FestDPO outperforms preference optimization baselines in both win rates against the base models and human evaluation scores. In protein backbone generation, it achieves a higher $β$-sheet fraction and better structural designability than the baselines.

CommentsPreprint. Under Review

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑