arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向文本到图像扩散模型的几何感知偏好优化

Geometry-Aware Preference Optimization for Text-to-Image Diffusion Models

Lei Wang, Zhen Wang, Yuexiang Xie, Yaliang Li

arXiv 2610.04980首次发表:更新:

发表机构

Sun Yat-Sen University; Alibaba Group(中山大学; 阿里巴巴集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对文本到图像扩散模型的偏好对齐,提出各向异性几何感知偏好优化(APO),利用参考模型的几何信息调整优化方向,提升图像质量与多样性,平均胜率超60%。

AI 中文摘要

偏好对齐已成为文本到图像扩散模型的标准实践。直接偏好优化(DPO)通过消除显式奖励建模简化了这一过程。其扩散变体 Diffusion-DPO 已成为广泛采用的基线方法。Diffusion-DPO 本质上鼓励偏好样本的似然性,同时抑制非偏好样本。在本文中,我们从流形假设的角度重新审视扩散模型的 DPO 风格对齐方法。在这种观点下,自然图像集中在嵌入高维环境空间的低维流形附近,而 DPO 直接在全空间中优化偏好分布,未考虑这种几何结构。这导致了优化动态中的不匹配:它抑制了保持几何的切向更新,同时未能充分限制危险的法向更新。这种不匹配逐渐降低了图像质量和多样性。为解决此问题,我们提出了各向异性几何感知偏好优化(APO),该方法将预测误差的均匀欧几里得处理替换为源自参考模型的几何感知各向异性度量。具体而言,APO 在参考去噪函数高度敏感的方向上自适应地加强正则化,同时在允许安全语义调整的方向上放松约束。这根据局部流形几何重新校准偏好优化,并保持原始流形结构。实验表明,APO 在多种基准测试中取得了强劲性能,平均胜率超过 60%,优于各种现有对齐方法。它所需的训练步骤显著少于先前方法,并在整个训练过程中保持了生成多样性。

英文摘要

Preference alignment has become a standard practice for text-to-image diffusion models. Direct Preference Optimization (DPO) simplifies this process by eliminating explicit reward modeling. Its diffusion variant, Diffusion-DPO, has become a widely adopted baseline. Diffusion-DPO essentially encourages the likelihood of preferred samples while suppressing dispreferred ones. In this paper, we revisit DPO-style alignment methods for diffusion models from the perspective of the manifold hypothesis. Under this view, natural images concentrate near a low-dimensional manifold embedded in the high-dimensional ambient space, whereas DPO directly optimizes preference distributions in the full space without accounting for this geometric structure. This creates a mismatch in the optimization dynamics: it suppresses geometry-preserving tangential updates, while insufficiently restricting hazardous normal-direction updates. This mismatch gradually degrades image quality and diversity. To address this issue, we propose Anisotropic Geometry-Aware Preference Optimization (APO), which replaces the uniform Euclidean treatment of prediction errors with a geometry-aware anisotropic metric derived from the reference model. Concretely, APO adaptively strengthens regularization in directions where the reference denoising function is highly sensitive, while relaxing constraints in directions that permit safe semantic adjustment. This recalibrates preference optimization according to the local manifold geometry, and maintains the original manifold structure. Experiments show that APO achieves strong performance and an average win rate exceeding 60\% against various existing alignment methods across diverse benchmarks. It requires significantly fewer training steps than prior methods, and preserves generation diversity throughout training.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑