arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

少即是多:视觉概念融合中正空间与负空间的平衡

Less Is More: Balancing Positive and Negative Space in Visual Concept Blending

Shishi Xiao, Adam J. Coscia, David H. Laidlaw

arXiv 2609.00476首次发表:更新:

发表机构

Brown University; Georgia Institute of Technology(布朗大学; 佐治亚理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出一种自动视觉概念融合流程,结合视觉语言模型与几何约束识别区域,采用混合像素-矢量方法生成融合构图,经评估在表现力等方面优于基线,可用于可控图像与信息图生成。

AI 中文摘要

平面设计师常通过融合视觉概念在单张图像中传递多重意涵,利用正空间与负空间实现平衡、强调效果与美学吸引力。尽管计算方法已开始支持自动概念融合,但大多忽略了设计中的空间构图作用。为解决这一差距,本文提出一种在融合过程中明确应用正空间与负空间的自动流程。该方法首先结合视觉语言模型的语义推理与真实案例的几何约束,识别出合理的概念整合区域;基于这些区域,系统采用混合像素-矢量流程生成融合构图:基于扩散的图像修复生成快速粗略初始化,随后通过基于点级的矢量优化进行细化,以确保结构连贯性与语义表达平衡。多模态智能体作为规划器与评估器协调该流程,支持迭代改进与可解释控制。通过基线对比与用户研究评估,本文方法因有效利用正空间与负空间,展现出更强的表现力、创造力与概念可识别性,还在可控图像与信息图生成等多样应用中验证了泛化性。

英文摘要

Graphic designers often blend visual concepts to communicate multiple ideas within a single image, leveraging positive and negative space to create balance, emphasis, and aesthetic appeal. While computational methods have begun to support automatic concept blending, they largely overlook the role of spatial composition in the design. To address this gap, we present an automatic pipeline that explicitly applies positive and negative space throughout the blending process. Our approach first identifies plausible regions for concept integration by combining semantic reasoning from vision-language models with geometric constraints derived from real-world examples. Conditioned on these regions, the system generates blended compositions using a hybrid pixel-vector pipeline: diffusion-based inpainting produces a fast, coarse initialization, which is then refined through vector-based optimization at the point level to ensure structural coherence and balanced semantic expression. A multimodal agent orchestrates this process as a planner and evaluator, enabling iterative improvement and interpretable control. Through an evaluation using both baseline comparisons and a user study, we demonstrate greater expressiveness, creativity, and concept recognizability by effectively leveraging positive and negative space. We further demonstrate the generalizability of our approach across diverse applications, including controllable image and infographic generation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑