StructGen:通过结构化上下文建模消除多参考图像生成中的歧义
StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling
浏览论文内容
中文总结 AI 辅助
研究多参考图像生成问题,现有方法因自然语言指令缺陷致效果不佳。提出StructGen,用结构化格式编码参考图像,构建数据集、训练框架和基准,实验证明其在语义对齐和生成一致性上优于现有方法。
中文摘要 AI 辅助
多参考图像生成旨在根据文本指令整合多个参考图像的属性来合成图像。随着参考数量增加,该任务需要复杂语义理解。现有仅依赖自然语言指令的方法常无法精确捕捉复杂意图,导致语义错位和生成不一致。我们发现自然语言指令冗长模糊以及高质量多参考数据稀缺是限制因素。为此提出StructGen,采用结构化、类似字典的格式编码多参考图像,明确生成意图。我们构建结构化数据集、开发训练框架及基准。实验表明StructGen在语义对齐和详细参考生成一致性上优于现有方法。
英文摘要
Multi-reference image generation aims to synthesize images by integrating attributes from multiple reference images under textual instructions. As the number of references increases, the task necessitates complex semantic comprehension, such as correctly associating attributes with the intended subjects and planing out coherent spatial arrangement between subjects and their environments. Existing approaches, which rely solely on natural language instruction, often fail to capture these complex intentions precisely, leading to semantic misalignment and inconsistent generation. We identify two key factors behind these limitations: natural language instructions are often verbose and ambiguous, and high-quality multi-reference data is scarce. To address these issues, we propose StructGen, which employs a structured, dictionary-like format to encode multiple reference images, thereby enabling explicit and unambiguous specification of generation intentions. To support this design, we construct a structured dataset based on high-quality real images and develop a corresponding training framework, along with a dedicated benchmark for challenging multi-reference scenarios. Extensive experiments on both public benchmarks and our proposed benchmark demonstrate that StructGen consistently outperforms existing methods on both semantic alignment and detailed reference-generation consistency, especially under complex instructions with multiple references. The code is available at https://jianingpeng0382.github.io/StructGen/
发表机构
- Institute of Information Science, Beijing Jiaotong University(北京交通大学信息科学研究所)
- Institute of Big Data, Fudan University(复旦大学大数据研究院)
- Visual Intelligence + X International Joint Laboratory of the Ministry of Education(教育部视觉智能+X国际联合实验室)
- Beijing Academy of Artificial Intelligence(北京人工智能研究院)
- MT Lab, Meitu Inc(美图公司MT实验室)
机构由 AI 辅助整理,请以论文原文为准。