基于声音的多人三维姿态估计
Sound-based Multi-Person 3D Pose Estimation
查看机构详情
- Keio University(庆应义塾大学)
- Tokyo University of Science(东京理科大学)
- NTT, Inc.(日本电信电话株式会社)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
本研究首次尝试仅从声学信号估计多人三维姿态,提出SoundMHPE编码器-解码器框架,构建含43.2万帧的AMP数据集,实验显示其性能优于基线模型。
中文摘要 AI 辅助
我们能否仅利用声音恢复多人的三维姿态?本文首次尝试仅从声学信号估计多人三维姿态。利用声学信号估计多人姿态固有地具有挑战性,因为存在依赖于运动的信号变化的叠加。与单人场景不同,多主体的存在会导致声学特征重叠,难以将特定信号变化归因于个体的姿态。此外,人与人之间的反射会引入复杂的传播延迟,模糊了时间上的运动-声学关系,进一步加剧了复杂性。为解决这些问题,我们提出SoundMHPE(基于声音的多人人体姿态估计器),这是一种新型编码器-解码器框架,包含两个关键组件:第一,声学多尺度编码器,捕获多样的时间和细粒度频率特征,以从复杂的重叠信号中分离出细微的声学特征;第二,时间姿态解码器,采用注意力机制在连续帧中解耦多人信息,通过联合考虑时间动态和人与人之间的依赖关系,该组件能精确重建逐帧的个体姿态。为验证我们的方法,我们构建了6小时的声学多人姿态(AMP)数据集,包含432,000帧同步的多人姿态与声学数据,实验表明我们的SoundMHPE性能优于基线模型。项目页面:this https URL
英文摘要
Can we recover the 3D poses of multiple people using only sound? This paper presents the first attempt to estimate multi-person 3D poses solely from acoustic signals. Estimating the poses of multiple individuals using acoustic signals is inherently challenging due to the superposition of motion-dependent signal variations. Unlike single-person scenarios, the presence of multiple subjects leads to overlapping acoustic signatures, making it difficult to attribute specific signal changes to an individual's pose. Furthermore, the complexity is compounded by inter-person reflections, which introduce intricate propagation delays that obscure the temporal motion-acoustic relationship. To address these issues, we propose SoundMHPE (Sound-based Multi-person Human Pose Estimator), a novel encoder-decoder framework consisting of two key components. First, the Acoustic Multi-scale Encoder captures diverse temporal and fine-grained frequency features to isolate subtle acoustic signatures from complex, overlapping signals. Second, the Temporal Pose Decoder employs an attention mechanism to disentangle multi-person information across successive frames. By jointly accounting for temporal dynamics and inter-person dependencies, this component precisely reconstructs frame-wise individual poses. To validate our approach, we constructed the 6-hour Acoustic Multi-person Pose (AMP) dataset consisting of 432K synchronized frames of multi-person pose and acoustic data, and demonstrated that our SoundMHPE outperforms baseline models. Project page: https://oumi03.github.io/sound-mhpe/