发表机构
School of Computer Science and Engineering, South China University of Technology; School of Future Technology, South China University of Technology(华南理工大学计算机科学与工程学院; 华南理工大学未来技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出InstancePin框架,通过坐标锚定与实例感知适配器等技术,解决布局到图像扩散模型的实例纠缠问题,在Cityscapes数据集上提升了图像保真度与语义一致性。
AI 中文摘要
布局到图像扩散模型通过以类别级分割图为条件生成图像,已实现出色的语义可控性,但这种类别对齐的控制不一定具备实例可寻址性:同一类别的多个邻近对象常被视为共享语义区域,导致边界模糊、外观平均化及实例间特征混淆。这一局限在城市场景合成中尤为明显,其中小型且密集的行人或车辆需要细粒度的实例分离,同时保持全局场景一致性。本文提出InstancePin,一种实例可寻址的布局到图像扩散框架,通过显式坐标锚定固定每个对象实例。InstancePin不直接将实例掩码注入预训练的骨干网络,而是引入独立的实例感知适配器,在保留类别级生成先验的同时学习实例特定的空间控制。对于每个实例,其中心坐标经傅里叶特征编码后投影为坐标令牌,该令牌作为空间锚定,通过坐标锚定注意力查询潜在图像特征。为使这些锚定具备空间意义,我们进一步用实例区域监督坐标注意力图,鼓励每个坐标令牌激活其对应的对象区域。最后,实例掩码引导的融合模块将预训练骨干网络特征路由至非实例区域,适配器特征路由至实例区域,实现局部实例细化且不牺牲全局语义保真度。在Cityscapes上的大量实验表明,InstancePin缓解了密集布局中的实例纠缠,提升了图像保真度与语义一致性。
英文摘要
Layout-to-image diffusion models have achieved impressive semantic controllability by conditioning generation on category-level segmentation maps. However, such category-aligned control is not necessarily instance-addressable: multiple nearby objects from the same category are often treated as a shared semantic region, leading to ambiguous boundaries, averaged appearances, and feature confusion among instances. This limitation is particularly evident in urban scene synthesis, where small and crowded pedestrians or vehicles require fine-grained instance separation while preserving global scene consistency. In this paper, we propose InstancePin, an instance-addressable layout-to-image diffusion framework that pins each object instance with an explicit coordinate anchor. Instead of directly injecting instance masks into the pretrained backbone, InstancePin introduces an independent instance-aware adapter to preserve the category-level generation prior while learning instance-specific spatial control. For each instance, its center coordinate is encoded with Fourier features and projected into a coordinate token, which serves as a spatial anchor queried by latent image features through coordinate pinning attention. To make these anchors spatially meaningful, we further supervise the coordinate attention maps with instance regions, encouraging each coordinate token to activate its corresponding object area. Finally, an instance-mask guided fusion module routes pretrained backbone features to non-instance regions and adapter features to instance regions, enabling local instance refinement without sacrificing global semantic fidelity. Extensive experiments on Cityscapes demonstrate that InstancePin mitigates instance entanglement in dense layouts and improves both image fidelity and semantic consistency.
CommentsAccepted to PRCV 2026