arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

X-MULTI:基于VLM的成像因子解耦用于因子感知图像合成

X-MULTI: VLM-based Imaging Factor Disentanglement for Factor-Aware Image Synthesis

Sonali Godavarthy, Matthias Neuwirth-Trapp, Tim-Felix Faasch, Maarten Bieshaar, Michael Moeller, Kristof Van Laerhoven, Danda Pani Paudel

arXiv 2608.24563首次发表:更新:

发表机构

University of Siegen; ETH Zürich; Bosch Research; INSAIT; Sofia University “St. Kliment Ohridski”(锡根大学; 苏黎世联邦理工学院; 博世研究中心; INSAIT(保加利亚的智能与数据科学研究所); 索非亚大学“圣·克利门特·奥赫里德斯基”)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对文本到图像生成的成像因子解耦问题,提出基于VLM的X-MULTI方法及改进的I-FAA指标,提升了新颖因子组合的因子对齐效果与评估的稳健性。

AI 中文摘要

文本到图像生成中的成像因子解耦旨在独立控制图像采集属性,如相机镜头类型、传感器类型、视角和领域,以实现组合泛化,使模型能合成训练数据中未出现的新颖因子组合,例如将训练数据中从未出现过的鱼眼镜头与事件传感器配对。近期研究MULTI引入了可学习的、因子特定的嵌入来解耦成像因子,并提出因子对齐准确率(FAA)指标评估解耦质量。我们识别并解决了两个独立的局限性:其一,MULTI的像素级重建目标仅在观测到的成像因子组合上监督模型,未对新颖组合提供直接训练信号,因此我们提出X-MULTI,使用预训练的视觉语言模型(VLM)监督训练期间合成的新颖因子组合;其二,我们发现FAA指标存在严重的跨因子相关性泄漏,会错误表征真实解耦质量,因此我们提出改进版FAA(I-FAA),采用因子特定的增强策略打破这些相关性,实现更严格的评估。实验表明,与MULTI相比,X-MULTI在新颖组合上实现了更优的因子对齐;此外,我们证明FAA中的相关性泄漏会扭曲对真实因子解耦的评估,而I-FAA可减少该泄漏,因此能提供更稳健的因子对齐评估。

英文摘要

Imaging factor disentanglement in text-to-image generation aims to independently control image acquisition properties such as types of camera lenses, sensor types, viewpoints, and domains to enable combinatorial generalization. This should let the model synthesize novel factor combinations unobserved in the training data, such as pairing a fisheye lens with an event sensor never observed in training data. Recent work, MULTI, introduced learnable, factor-specific embeddings to disentangle imaging factors, along with the Factor Alignment Accuracy (FAA) metric to evaluate disentanglement quality. We identify and address two independent limitations. First, MULTI's pixel-level reconstruction objective supervises the model only on observed imaging factor combinations, providing no direct training signal for novel combinations. We therefore propose X-MULTI, which uses a pretrained vision-language model (VLM) to supervise novel factor combinations synthesized during training. Second, we show the FAA metric exhibits severe cross-factor correlation leakage, misrepresenting true disentanglement quality. We therefore propose Improved-FAA (I-FAA), which employs factor-specific augmentation strategies to break these correlations and enables more rigorous evaluation. Experiments demonstrate that X-MULTI achieves improved factor alignment on novel combinations compared to MULTI. Moreover, we show that correlation leakage in FAA distorts the evaluation of true factor disentanglement and I-FAA reduces this leakage and therefore provides a more robust assessment of factor alignment.

CommentsAccepted to the MUCG Workshop at ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑