arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PoseAdapter:面向复杂多目标场景的双流2.5D可控图像生成

PoseAdapter: Dual-Stream 2.5D Controllable Image Generation for Complex Multi-Object Scenes

Yufeng Chi, Huimin Ma, Fan Gao, Zhice Niu, Keqin Li, Jianmin Li

arXiv 2608.15583首次发表:更新:

发表机构

Tsinghua University; National University of Defense Technology(清华大学; 国防科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PoseAdapter是一种轻量2.5D可控图像生成框架,通过双流表示解决多目标场景属性泄漏问题,构建OrientLayout数据集,在空间精度等指标上优于现有最优方法。

AI 中文摘要

尽管文本到图像(T2I)扩散模型已取得显著成功,但多目标场景中的精确空间与方向控制仍是持续存在的挑战。现有方法要么依赖计算成本高昂的密集3D地图,要么存在严重的属性泄漏和“剪切粘贴”伪影。为解决这些局限,我们提出PoseAdapter——一种用于高保真2.5D可控图像生成的轻量框架。它不使用密集空间地图,而是通过高效的条件布局(单个目标描述、2D边界框和3D角度)建立精确的空间-角度锚点。为解决严格实例隔离与全局一致性之间的生成权衡,我们引入上下文感知双流表示:通过并行的掩码与非掩码路径将局部目标标记和富含关系的场景标记注入现代MM-DiT架构的视觉流,PoseAdapter在消除属性泄漏的同时保留自然的目标间关系与场景级一致性。为支撑该范式,我们构建了OrientLayout数据集,该数据集具有标准化2.5D标注和实例级解耦语义。大量实验表明,PoseAdapter在空间精度、方向准确度和多目标视觉保真度上优于现有最优基线。代码与数据集将在此URL提供。

英文摘要

While Text-to-Image (T2I) diffusion models have achieved remarkable success, precise spatial and orientational control in multi-object scenes remains a persistent challenge. Existing methods either rely on computationally expensive dense 3D maps or suffer from severe attribute leakage and "cut-and-paste" artifacts. To address these limitations, we propose PoseAdapter, a lightweight framework for high-fidelity 2.5D controllable image generation. Instead of dense spatial maps, it establishes precise spatial-angular anchors using an efficient condition layout: individual object captions, 2D bounding boxes, and 3D angles. To resolve the generative trade-off between strict instance isolation and global coherence, we introduce a Context-Aware Dual-Stream Representation. By injecting local object tokens and relation-enriched scene tokens into the visual stream of modern MM-DiT architectures via parallel masked and unmasked pathways, PoseAdapter eliminates attribute leakage while preserving natural inter-object relationships and scene-level coherence. To support this paradigm, we construct OrientLayout, a high-quality dataset featuring standardized 2.5D annotations and instance-level decoupled semantics. Extensive experiments demonstrate that PoseAdapter outperforms state-of-the-art baselines in spatial accuracy, orientational precision, and multi-object visual fidelity. Code and dataset will be available at https://github.com/cyf23/PoseAdapter.

Comments10 pages, 5 figures

DOI:10.1145/3767308.3835074

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑