发表机构
College of Intelligence and Computing, Tianjin University(天津大学智能与计算学部)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出一种带有人感知3D表示对比学习的前馈多视图多人重建方法,通过自顶向下范式和空间对比学习策略,在真实场景中实现了鲁棒高效的多人重建。
AI 中文摘要
多视图人体重建已在简化设置下得到广泛研究,但在非约束环境中实现鲁棒且高效的多人重建仍然具有挑战性。现有的自底向上方法通常依赖精确的相机标定和显式的跨视图匹配,因此在严重遮挡和歧义场景下表现不佳。我们提出一种新的自顶向下范式,该范式维护统一的、以实例为中心的人感知3D空间,通过跨模态对比学习实现相机标定、跨视图关联和人体重建的同步进行。多视图观测被提升并融合到该共享3D空间中,在该空间中,几何结构、视觉外观和以人为中心的语义线索在实例级别被联合编码。我们进一步引入空间对比学习策略,该策略对齐不同视图和模态中对应同一人体实例的3D特征,同时区分不同实例。这使得对应推理、语义聚合和实例区分能够在3D空间中自然执行,提高跨视图一致性和严重遮挡下的鲁棒性。最后,通过从实例级3D人体token回归SMPL参数,以前馈方式恢复结构化人体模型。大量实验表明,该方法在具有挑战性的真实场景中实现了鲁棒、准确且高效的多视图人体重建。
英文摘要
Multi-view human reconstruction has been extensively studied under simplified settings, yet robust and efficient multi-person reconstruction in unconstrained environments remains challenging. Existing bottom-up methods often rely on accurate camera calibration and explicit cross-view matching, and therefore struggle with severe occlusions and ambiguities. We propose a new top-down paradigm that maintains a unified, instance-centric human-aware 3D space, enabling simultaneous camera calibration, cross-view association, and human reconstruction via cross-modal contrastive learning. Observations from multiple views are lifted and fused into this shared 3D space, where geometric structure, visual appearance, and human-centric semantic cues are jointly encoded at the instance level. We further introduce a spatial contrastive learning strategy that aligns 3D features corresponding to the same human instance across different views and modalities while separating different instances. This enables correspondence reasoning, semantic aggregation, and instance discrimination to be performed natively in 3D, improving cross-view consistency and robustness under severe occlusions. Finally, structured human body models are recovered in a feed-forward manner by regressing SMPL parameters from instance-level 3D human tokens. Extensive experiments demonstrate robust, accurate, and efficient multi-view human reconstruction in challenging real-world scenarios.
CommentsPublished in International Journal of Computer Vision (IJCV)
Journal refInternational Journal of Computer Vision 134, 414 (2026)
DOI:10.1007/s11263-026-03000-0