超越原子布局:基于视觉-语言模型的组合式设计理解
Beyond Atomic Layouts: Compositional Design Understanding with Vision-Language Models
浏览论文内容
中文总结 AI 辅助
本文针对组合式布局理解任务,提出后训练范式MASON,利用约2万布局的CoDeLayout数据集,使Qwen2.5-VL 7B准确率达91.66%,优于基线及全数据微调。
中文摘要 AI 辅助
布局理解即对元素组织方式的解读,对文档分析、用户界面(UI)创建及平面设计至关重要。尽管近期的视觉-语言模型(VLMs)擅长解读由独立元素构成的原子布局,但它们在处理组合式布局时存在困难,这类布局需要对分层多层结构内视觉上纠缠的元素进行推理。本文提出一项新任务——组合式布局理解,并发布CoDeLayout,这是一个包含约2万个真实世界多层布局的视觉问答(VQA)数据集,标注了组合式元素对及设计意图。通过对CoDeLayout的实证分析,我们明确现有VLMs面临两个关键挑战:文本元数据与视觉内容间的语义漂移,以及分层元素间关系的结构歧义。为应对这些挑战,我们提出MASON,一种整合多模态对齐(MA)与结构感知(SP)的后训练范式。MA通过将元数据定义的元素与其视觉对应项关联,增强元素解读,缓解语义漂移;SP则对分层感知的元素间空间关系建模,以提升分层理解并减少结构歧义。实验显示现有VLMs存在显著差距:最强基线GPT-o3仅达到79.68%的准确率,而采用MASON的Qwen2.5-VL 7B达到91.66%。值得注意的是,MASON仅使用30%的训练数据就超越了全数据直接微调,且随着数据增加扩展性更好。
英文摘要
Layout understanding, or the interpretation of element organization, is essential for document analysis, user interface (UI) creation, and graphic design. While recent vision-language models (VLMs) excel at interpreting atomic layouts composed of independent elements, they struggle with compositional layouts that require reasoning over visually entangled elements within hierarchical multi-layer structures. In this paper, we introduce a new task, compositional layout understanding, and present CoDeLayout, a VQA dataset of ~20K real-world multi-layer layouts annotated with compositional element pairs and design intent. Through empirical analysis on CoDeLayout, we identify two key challenges for existing VLMs: semantic drift between textual metadata and visual content, and structural ambiguity in hierarchical inter-element relationships. To address these challenges, we propose MASON, a post-training paradigm that integrates multimodal alignment (MA) and structural perception (SP). MA enhances element interpretation by grounding metadata-defined elements to their visual counterparts, mitigating semantic drift, while SP models layer-aware inter-element spatial relationships to improve hierarchical understanding and reduce structural ambiguity. Experiments reveal substantial gaps in existing VLMs: even the strongest baseline, GPT-o3, achieves only 79.68% accuracy, whereas Qwen2.5-VL 7B with MASON reaches 91.66%. Notably, MASON surpasses full-data Direct Finetune using only 30% of the training data and scales better with additional data.
发表机构
- Northeastern University(东北大学)
- Adobe Research(奥多比研究院)
机构由 AI 辅助整理,请以论文原文为准。