arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OccluRank:仅添加序数秩的可控遮挡感知布局到图像生成

OccluRank: Controllable Occlusion-Aware Layout-to-Image Generation by Adding Just an Ordinal Rank

Wenyang Hong, Yuan Wang, Yanbin Hao, Lanqing Xue, Ke Wang, Xiang Wang, Kuien Liu, Richang Hong

arXiv 2608.20932首次发表:更新:

AI 中文总结

提出OccluRank框架,仅添加序数秩实现可控遮挡感知布局到图像生成,构建相关数据集与基准,实验表明其能可靠保留实例、遵循布局并实现期望遮挡关系,性能相当。

AI 中文摘要

布局到图像生成通过边界框布局实现显式空间控制,但边界框仅能指定实例位置,无法表示它们的遮挡顺序。现有方法可能依赖额外几何条件、采用复杂推理流程,或聚合独立构建的实例表示却未显式建模其依赖遮挡的交互。我们提出OccluRank,一种简单且可控的遮挡感知布局到图像框架,仅用一个序数秩扩充每个边界框。OccluRank通过轻量的基于秩的条件编码用户指定的遮挡顺序,并引入顺序感知实例交互(Order-aware Instance Interaction,OII)模块,在聚合前联合更新经秩条件处理的实例表示。这使得指定顺序能引导遮挡实例间的信息交换,无需额外几何输入或专用推理时优化。我们进一步构建OccluLayout,这是一个合成训练数据集,其遮挡顺序和无模态标注直接源自已知场景几何,而非用辅助预测模型从部分遮挡图像中估计。为全面评估,我们引入OccluLayout-Bench,该基准使用多个多模态大语言模型评估器评估实例存在性、空间布局、属性及遮挡顺序,同时用FID评估整体图像质量。实验表明,OccluRank能更可靠地保留目标实例、遵循指定布局并实现期望的遮挡关系,同时保持相当的属性一致性和整体图像质量。

英文摘要

Layout-to-image generation enables explicit spatial control through bounding-box layouts, yet bounding boxes specify only instance locations and cannot represent their occlusion order. Existing methods may rely on additional geometric conditions, employ complex inference procedures, or aggregate independently constructed instance representations without explicitly modeling their occlusion-dependent interactions. We propose OccluRank, a simple and controllable occlusion-aware layout-to-image framework that augments each bounding box with only one ordinal rank. OccluRank encodes the user-specified occlusion order through lightweight rank-based conditioning and introduces an Order-aware Instance Interaction (OII) module to jointly update rank-conditioned instance representations before aggregation. This allows the specified order to guide information exchange among occluding instances without additional geometric inputs or specialized inference-time optimization. We further construct OccluLayout, a synthetic training dataset whose occlusion order and amodal annotations are derived directly from known scene geometry rather than estimated from partially occluded images using auxiliary prediction models. For comprehensive evaluation, we introduce OccluLayout-Bench, which uses multiple multimodal large language model evaluators to assess instance presence, spatial layout, attributes, and occlusion order, together with FID for overall image quality. Experiments show that OccluRank more reliably preserves target instances, follows specified layouts, and realizes desired occlusion relationships while maintaining comparable attribute consistency and overall image quality.

Comments16 pages, 7 figures. Code: https://github.com/Wenyang-hong/OccluRank

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑