发表机构
UC San Diego; Lambda, Inc.(加州大学圣迭戈分校; Lambda 公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有布局到图像数据集缺乏密集复杂交互的问题,提出OverLay++大规模数据集,含50万图像、平均每图6.6对象,标注密度和对象描述显著提升,训练SOTA方法带来持续改进和更快收敛。
AI 中文摘要
布局到图像生成在空间和对象级控制方面取得了实质性进展。然而,现有方法在处理包含许多重叠和交互对象的复杂场景时仍然存在困难。我们认为训练数据是一个特别的瓶颈:现有数据集缺乏具有密集、复杂对象交互的示例。为解决这一差距,我们引入了OverLay++,一个具有结构复杂场景的大规模布局到图像数据集。OverLay++包含约50万张图像,每张图像平均有6.6个对象,在标注密度上超过现有数据集1.67倍。除了标注密度,OverLay++还提供了丰富的语义细节,其对象描述比当前数据集长六倍以上。我们的数据集生成流程简单,并生成具有丰富逐对象描述的密集重叠对象标注。在多个基准测试中,基于OverLay++数据集训练的最先进布局到图像方法表现出持续改进和更快的收敛速度,证明了密集、重叠感知和描述丰富的监督对于可控图像生成的重要性。
英文摘要
Layout-to-Image generation has made substantial progress in spatial and object-level control. However, existing methods still struggle with complex scenes containing many overlapping and interacting objects. We argue that training data is a particular bottleneck: existing datasets lack examples with dense, complex object interactions. To address this gap, we introduce OverLay++, a large-scale Layout-to-Image dataset with structurally complex scenes. OverLay++ contains approximately 500K images with an average of 6.6 objects per image, exceeding existing datasets by 1.67 times in annotation density. Beyond annotation density, OverLay++ provides rich semantic detail with object captions over six times longer than in current datasets. Our dataset generation pipeline is simple and produces dense, overlapping object annotations with rich per-object captions. Across multiple benchmarks, state-of-the-art Layout-to-Image methods trained on the OverLay++ dataset show consistent improvement and faster convergence, demonstrating the importance of dense, overlap-aware, and caption-rich supervision for controllable image generation.
CommentsAccepted at NeurIPS 2026, Evaluations & Datasets Track. Project website: https://mlpc-ucsd.github.io/OverLayPP . Dataset: https://huggingface.co/datasets/mlpcucsd/OverLayPP