arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

UFO:多模态图像生成中全条件对齐的评估链

UFO: Chain-of-Evaluation for Omni-Condition Alignment in Multi-Modal Image Generation

Danning Zhang, Yijing Lin, Shuhan Zhuang, Mengqi Huang, Shaojin Wu, Shancheng Fang, Zhendong Mao

arXiv 2609.12397首次发表:更新:

AI 中文总结

针对多模态图像生成中现有评估方法孤立检查各模态条件的问题,提出统一框架UFO,通过原子化评估链分解全条件对齐并分类验证,实现与人类偏好最高相关性,平均提升15.25%,并构建基准UFO-Bench。

AI 中文摘要

多模态图像生成,特别是主体驱动的定制化,近年来受到越来越多的关注。尽管生成模型发展迅速,但其评估方法仍明显滞后。现有方法,无论是基于嵌入的还是基于多模态大语言模型(MLLM)的,都孤立地评估与每种模态条件的对齐,这与多模态图像生成的同时条件对齐目标相矛盾,导致与人类判断的一致性较差。为应对这一挑战,我们提出了UFO,这是首个用于全条件对齐同时评估的统一框架。具体而言,UFO引入了一种新颖的原子化评估链范式,即首先将全条件对齐分解为一系列细粒度、解耦的原子评估单元(AEU),将它们分类为不同的模态相关类别,然后采用通用或专用的功能调用对不同AEU类型进行准确验证。实验结果表明,UFO与人类评估偏好的相关性最高,平均提升了15.25%。此外,我们提出了UFO-Bench,一个专门的基准,旨在在文本和视觉条件的多样相互交互下全面评估现有定制化模型的性能。

英文摘要

Multi-modal image generation, particularly subject-driven customization, has garnered growing attention in recent years. Despite the rapid advancement of generative models, their evaluation remains largely lagging. Existing methods, whether embedding-based or Multi-modal Large Language Model (MLLM)-based, evaluate alignment with each modal condition in isolation, which contradicts the simultaneous condition alignment objective of multi-modal image generation, leading to poor consistency with human judgments. To address this challenge, we propose UFO, the first unified framework for omni-condition alignment simultaneous evaluation. Specifically, UFO introduces a novel Atomized Chain-of-Evaluation paradigm, i.e., it first decomposes omni-condition alignment into a sequential chain of fine-grained, disentangled Atomic Evaluation Units (AEUs), categorizes them into distinct modality-relevance classes, and then employs general or dedicated functional calls for accurate verification of different AEU types. Experimental results demonstrate that UFO achieves the highest correlation with human evaluation preferences, delivering an average improvement of 15.25%. Furthermore, we present UFO-Bench, a dedicated benchmark designed to holistically evaluate the performance of existing customization models under the diverse mutual interactions of textual and visual conditions.

Comments13pages, 6 figures, accepted at the Forty-Third International Conference on Machine Learning (ICML 2026)

Journal refProceedings of the Forty-Third International Conference on Machine Learning, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑