发表机构
Beijing University of Posts and Telecommunications(北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有文本到全景生成方法难以对齐对象级方向描述的问题,提出以对象为中心的可控框架PanoCtrl,构建相关数据集并实现最优性能。
AI 中文摘要
全景图像生成对于虚拟现实、增强现实和3D内容创建等沉浸式应用愈发重要。与透视图像不同,全景图像表示以观察者为中心的360°周围空间,其中左、右、前、后等方向表达在空间理解中起核心作用。然而,现有的文本到全景生成方法大多依赖隐式空间推理,常无法在球面全景场景中忠实地锚定对象级别的方向描述。引入显式布局是一种直接替代方案,但需要手动指定空间条件会降低语言交互的灵活性,且无法直接解决自我中心方向语言与全景图像空间之间的错位问题。为解决该问题,我们提出PanoCtrl,一种用于可控文本到全景生成的以对象为中心的框架。我们的方法通过将文本描述转换为结构化的对象级球面条件并将其融入扩散过程,明确连接自然语言与球面全景空间。具体而言,我们引入PanoParse,一种文本条件解析器,可预测对象语义和球面边界视场(BFoV)参数;还引入PanoControl,通过面向对象的注意力机制和空间残差增强,将对象级语义和空间指导注入扩散Transformer。为支持该任务,我们构建了PanoGround数据集,包含用于可控全景生成的对象级球面标注和多样的方向描述。大量实验表明,PanoCtrl在空间对齐和图像质量方面均达到了当前最优性能。
英文摘要
Panoramic image generation is increasingly important for immersive applications such as virtual reality, augmented reality, and 3D content creation. Unlike perspective images, panoramic images represent a viewer-centered $360^\circ$ surrounding space, where directional expressions such as left, right, front, and behind play a central role in spatial understanding. However, existing text-to-panorama methods largely rely on implicit spatial reasoning and often fail to faithfully ground object-level directional descriptions in spherical panoramic scenes. A straightforward alternative is to introduce explicit layouts, but requiring manually specified spatial conditions reduces the flexibility of language-based interaction and does not directly resolve the misalignment between egocentric directional language and panoramic image space. To address this issue, we propose PanoCtrl, an object-centric framework for controllable text-to-panorama generation. Our method explicitly bridges natural language and spherical panoramic space by converting textual descriptions into structured object-level spherical conditions and integrating them into the diffusion process. Specifically, we introduce PanoParse, a text-conditioned parser that predicts object semantics and spherical bounding field-of-view (BFoV) parameters, and \textbf{PanoControl}, which injects object-level semantic and spatial guidance into the diffusion transformer through object-aware attention and spatial residual enhancement. To support this task, we construct PanoGround, a dataset with object-level spherical annotations and diverse directional descriptions for controllable panoramic generation. Extensive experiments demonstrate that PanoCtrl achieves state-of-the-art performance in both spatial alignment and image quality.