arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大场景下有遮挡的多视角多人人体网格恢复

Multiview Multi-Person Human Mesh Recovery Under Large Scenes with Occlusions

Qi Zhang, Tao Yu, Jiechao He, Antoni B. Chan, Hui Huang

arXiv 2607.24302首次发表:更新:

发表机构

Guangdong Provincial Key Laboratory of Visual Media and Multidimensional Intelligence, College of Computer Science and Software Engineering, Shenzhen University; City University of Hong Kong(广东省视觉媒体与多维智能重点实验室,深圳大学计算机科学与软件学院; 香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对现有人体网格恢复基准和方法在大场景及严重遮挡下的不足,引入MVMP-HMR基准,提出MVMP-HMR模型,通过融合多视角特征、利用骨盆关节提取查询、交叉注意力整合及引入新损失,提升了大场景严重遮挡下的人体网格恢复效果。

AI 中文摘要

人体网格恢复旨在从图像中恢复3D人体网格。现有的大多数人体网格恢复基准和方法要么专注于单视角多人重建,要么专注于多视角单人重建,场景规模和人物数量有限,不足以应对大场景和严重遮挡的实际应用。为此,我们引入了一个用于多视角多人人体网格恢复的大规模合成基准MVMP-HMR,它包含15个复杂场景,有多达50个相机视角和30个交互人物,具有大空间覆盖和严重遮挡,增加了人体网格恢复难度。基于此基准,我们进一步提出了多视角多人全身人体网格恢复模型MVMP-HMR模型。该模型先将多视角特征融合成场景级3D特征体,再利用3D姿态估计网络预测的骨盆关节从3D特征体中提取特定人物查询,这些查询与3D特征体交叉注意力并整合以解码每个人的3D网格。此外,我们引入了两个新损失——方向损失和3D关节密度损失——以减轻严重遮挡下的方向和姿态模糊性。实验表明,现有最先进的人体网格恢复方法在MVMP-HMR基准上表现不佳,而我们的方法在大场景严重遮挡下始终优于先前的方法。

英文摘要

Human mesh recovery (HMR) aims to recover 3D human meshes from images. Most existing HMR benchmarks and methods focus on either multi-person reconstruction from a single view or single-person reconstruction from multiple views, where the number of subjects and the scene scale are relatively limited. Such settings are insufficient for real-world applications with large scenes and severe inter-person occlusions. To address this limitation, we introduce a large-scale synthetic benchmark for multiview multi-person HMR, termed MVMP-HMR. The proposed dataset contains 15 complex scenes with up to 50 camera views and 30 interacting persons, featuring large spatial coverage and severe occlusions, which significantly increases the difficulty of human mesh recovery. Based on this benchmark, we further propose a multiview multi-person whole-body human mesh recovery model, referred to as MVMP-HMR model. The model first fuses multiview features into a scene-level 3D feature volume, and then leverages pelvis joints predicted by a 3D pose estimation network to extract person-specific queries from the 3D feature volume. These human queries are cross-attended with the 3D feature volume and integrated to decode each person's 3D mesh. Moreover, we introduce two novel losses--the orientation loss and the 3D joint density loss--to alleviate orientation and pose ambiguities under severe occlusions. Experiments demonstrate that existing state-of-the-art HMR methods struggle on the proposed MVMP-HMR benchmark, while our method consistently outperforms prior SOTAs in large-scale scenes with severe occlusions.

Comments10 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑