发表机构
Shenzhen University(深圳大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多视角扩散模型在城市街道场景中跨视角一致性差的问题,提出StreetDiff框架,通过全景-透视协同设计和全景对齐模块显式施加跨视角对齐约束,并构建Street360数据集,显著提升结构一致性与视觉保真度。
AI 中文摘要
多视角扩散模型在具有强几何先验和稀疏语义的场景(如室内房间或简单室外环境,如田野、庭院)中表现出色。然而,在相机旋转情况下,它们往往难以维持跨视角一致性,尤其是在结构复杂的城市环境中。由于缺乏对跨视角球面对应关系的显式建模,现有方法容易产生物体重复、结构扭曲和布局不一致等问题。为解决这一局限,我们提出了StreetDiff,一种在去噪过程中显式强制跨视角对齐的多视角扩散框架。StreetDiff引入了全景-透视协同(Panorama--Perspective Synergy)设计,将全局布局推理与局部细节合成解耦,并整合了全景对齐模块(PAM),该模块建立基于球面投影的跨视角注意力约束。通过在不修改扩散主干的情况下注入结构化对齐约束,我们的框架在具有挑战性的城市街道场景生成任务中实现了鲁棒的跨视角一致性。此外,我们构建了Street360,一个大规模HDR多视角城市全景数据集。大量实验表明,与先前的多视角扩散生成方法相比,StreetDiff显著提高了结构一致性和视觉保真度。
英文摘要
Multi-view diffusion models have shown strong performance in scenes with strong geometric priors and sparse semantics, such as indoor rooms or simple outdoor environments (e.g., fields, courtyards). However, they often fail to maintain cross-view consistency under camera rotation, especially in structurally complex urban environments. Without explicit modeling of spherical correspondence across views, existing approaches tend to produce object duplication, structural distortion, and layout inconsistency. To address this limitation, we propose StreetDiff, a multi-view diffusion framework that explicitly enforces cross-view alignment during denoising. StreetDiff introduces a Panorama--Perspective Synergy design to decouple global layout reasoning from local detail synthesis, and incorporates a Panorama Alignment Module (PAM) that establishes spherical-projection-based attention constraints across views. By injecting structured alignment constraints without modifying the diffusion backbone, our framework achieves robust cross-view coherence in challenging urban street scene generation tasks. In addition, we construct Street360, a large-scale HDR multi-view urban panorama dataset. Extensive experiments demonstrate that StreetDiff significantly improves structural consistency and visual fidelity compared to prior multi-view diffusion generation methods.
Comments10 pages, 4 figures