发表机构
University of Chinese Academy of Sciences; Renmin University of China; Shanghai Theatre Academy; Institute of Automation, Chinese Academy of Sciences(中国科学院大学; 中国人民大学; 上海戏剧学院; 中国科学院自动化研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出FaV-A,一种自主多模态代理,通过渐进式流程分阶段生成正负空间构图,实验证明其优于直接零样本MLLM基线,能生成视觉连贯且语义对齐的图像。
AI 中文摘要
正负空间是视觉构图中的基本原理,支撑着视觉上连贯的形式和分层的语义关系。生成此类构图具有挑战性,因为它需要对共享共同边界的两个语义概念进行协调控制。尽管最近的文本到图像模型和多模态大语言模型(MLLMs)在图像生成和视觉理解方面取得了强劲性能,但正负空间生成仍然困难,尤其是在直接单次提示下。在这项工作中,我们提出了“形与空代理”(FaV-A),一种专为分阶段正负空间生成而设计的多模态代理。FaV-A遵循渐进式工作流程:它首先生成一个基础对象,然后分析其形状和空间结构以识别候选的负空间语义,最后为最终图像生成阶段生成构图指令。实验结果和消融分析表明,与直接零样本MLLM基线相比,FaV-A为生成视觉上连贯且语义对齐的正负空间构图提供了更有效的框架。
英文摘要
Positive and negative space is a fundamental principle in visual composition, supporting visually coherent forms and layered semantic relationships. Generating such compositions is challenging because it requires coordinated control over two semantic concepts that share a common boundary. Although recent text-to-image models and multimodal large language models (MLLMs) have achieved strong performance in image generation and visual understanding, positive-negative space generation remains difficult, particularly under direct single-pass prompting. In this work, we present the \textbf{F}orm \textbf{a}nd \textbf{V}oid \textbf{A}gent (\textbf{FaV-A}), a multimodal agent designed for staged positive-negative space generation. FaV-A follows a progressive workflow: it first generates a base object, then analyzes its shape and spatial structure to identify candidate negative-space semantics, and finally produces compositional instructions for the final image generation stage. Experimental results and ablation analyses suggest that FaV-A provides a more effective framework than direct zero-shot MLLM baselines for producing visually coherent and semantically aligned positive-negative space compositions.
CommentsCVPR Workshops AI4VA, 2026, Best Paper Award
Journal refProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2026, pp. 8987-8995