arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23302cs.CV

基于自由形式指令的时尚互补图像生成的接地

Grounding Free-Form Instructions for Fashion Complementary Image Generation

Matteo Attimonelli, Claudio Pomo, Alessandro De Bellis, Danilo Danese, Dietmar Jannach, Tommaso Di Noia

首次发表
浏览论文内容

中文总结 AI 辅助

针对现有时尚互补图像生成基准的缺陷,该研究提出基于自由形式指令的新设置,用StyleFlow模型实现该任务,其生成的服装符合指令且风格连贯,同时降低了架构复杂度和推理成本。

中文摘要 AI 辅助

时尚互补图像生成(CIG)旨在根据用户意图创建在风格上与种子单品匹配的服装,这是一个自然的多模态接地问题,模型必须在视觉上下文中解释语言。现有的CIG基准依赖于僵化的模板提示(例如“一条裙子的照片”),无法反映自然用户查询,且掩盖了模型在不同语言特异性水平下的行为。我们引入了带有自由形式指令的时尚互补图像生成,这是一种多模态语言接地设置,其中模型根据种子图像和自然语言指令生成兼容的服装。为此,我们用视觉语言模型生成的、经人工标注者验证的低、中、高特异性指令丰富了三个CIG基准。我们用StyleFlow(一个Rectified Flow Matching模型)实例化该任务,该模型在单个多模态Transformer中同时以种子图像和指令为条件。在图像质量指标、目录对齐分析、消融实验和人工评估中,StyleFlow始终生成符合指令且风格连贯的服装,同时相对于辅助模块方法降低了架构复杂度和推理成本。

英文摘要

Fashion complementary image generation (CIG) aims to create garments that stylistically match a seed item based on user intent, making it a natural multimodal grounding problem where models must interpret language in visual context. Existing CIG benchmarks rely on rigid template prompts (e.g., "a photo of a skirt"), failing to reflect natural user queries and obscuring model behavior across levels of linguistic specificity. We introduce fashion complementary image generation with free-form instructions, a multimodal language-grounding setting where a model generates a compatible garment from a seed image and a natural-language instruction. To this end, we enrich three CIG benchmarks with low-, medium-, and high-specificity instructions generated by a vision-language model and validated by human annotators. We instantiate the task with StyleFlow, a Rectified Flow Matching model that jointly conditions on the seed image and instruction within a single multimodal transformer. Across image quality metrics, catalog-alignment analysis, ablations, and human evaluation, StyleFlow consistently produces instruction-aligned and stylistically coherent garments while reducing architectural complexity and inference cost relative to auxiliary-module approaches.

发表机构

  • Politecnico di Bari(巴里理工大学)
  • Sapienza University of Rome(罗马大学)
  • University of Klagenfurt(克拉根福大学)

机构由 AI 辅助整理,请以论文原文为准。

↑