arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DiTailed:在文本-图像到图像流匹配模型中确保视觉对象一致性

DiTailed: Ensuring Visual Object Consistency in Text-Image-to-Image Flow Matching Models

Francesco Taioli, Daniel Coelho, Iaroslav Melekhov, Roberto Alcover-Couso, Jose Miguel Grande Saiz, Virginia Fernandez Arguedas, Artur Bekasov

arXiv 2607.12539首次发表:更新:

发表机构

Amazon; Faculty(亚马逊; 学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究文本引导图像编辑中生成模型视觉对象一致性问题,提出ABO-Edit数据集,发现图像编辑整流流模型条件嵌入空间特性,进而提出FlowMirror无参数辅助损失,无需架构改变提升了生成质量。

AI 中文摘要

尽管文本引导图像编辑取得显著进展,但生成模型常无法保持视觉对象一致性,即在整个编辑过程中保留主体关键属性。我们通过三项贡献解决此局限。首先,引入ABO-Edit数据集,专为研究对象一致性设计,含超12000个源图像、编辑提示及高质量目标图像三元组,有多视图覆盖和人工验证质量控制。其次,发现图像编辑整流流模型一个被忽视的属性:训练中未直接监督的条件嵌入空间,即使在高噪声水平下也编码最终生成图像的预测。第三,利用此发现提出FlowMirror,一种无参数辅助损失监督该条件嵌入空间。无需架构改变,我们的方法在多个指标上优于基线提高了生成质量。

英文摘要

Despite remarkable progress in text-guided image editing, generative models frequently fail to preserve visual object consistency, defined as the preservation of a subject's key attributes throughout the editing process. We address this limitation through three contributions. First, we introduce ABO-Edit, a dataset specifically designed to study object consistency, comprising over 12,000 triplets of source images, editing prompts, and high-quality target images rendered from artist-designed 3D assets, with multi-view coverage and human-verified quality control. Second, we uncover an overlooked property of image-editing rectified flow models: the conditioning embedding space, not directly supervised during training, encodes a prediction of the final generated image even at high noise levels. Third, exploiting this finding, we propose FlowMirror, a parameter-free auxiliary loss that supervises this conditioning embedding space. Without architectural changes, our method improves generation quality across several metrics over baselines.

CommentsAccepted to ECCV 2026. Project page: https://francescotaioli.github.io/DiTailed/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑