EgoGVAE:基于引导变分自编码器的自我身体网格重建
EgoGVAE: Ego-body Mesh Reconstruction via Guided Variational Autoencoder
- Konkuk University(建国大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对仅用头部姿态重建全身网格的难题,提出EgoGVAE方法,通过引导变分自编码器实现一步采样,推理速度比扩散方法快50倍以上,在基准数据集上提升了重建性能。
AI中文摘要:
我们解决仅从头部姿态恢复全身网格的问题,该任务对基于头戴设备或智能眼镜的各类应用至关重要。此任务的挑战在于仅根据单一关节(即头部)轨迹估算未观察到的身体部位姿态信息。已有研究开始采用头部条件生成模型,但这类方法因基于扩散的迭代过程而成本高、耗时长。作为替代方案,我们提出一种简单却新颖的方法,该方法利用引导网络的潜在空间,该网络被设计为以全身姿态为输入的变分自编码器。通过强制该引导网络与我们的头部到动作网络的潜在分布相似,从“引导”分布(即我们的头部到动作网络中学习到的分布)采样的潜在特征可被可靠解码,从而仅用头部姿态就能生成自然的全身姿态表示。所提方法的一个重要优势是,其一步采样方案的推理速度比基于扩散的方法快50倍以上。在基准数据集上的实验结果表明,所提方法有效提升了自我身体网格重建的性能。
英文摘要:
We address the problem of recovering the full-body mesh from only the head pose. This task has become essential for various applications based on head-mounted devices or smart glasses. The challenge of this task lies in estimating the pose information of unobserved body parts based solely on a single joint (i.e., head) trajectory. Several studies have begun to adopt head-conditioned generative models, however, such previous methods are costly and time-consuming due to the diffusion-based iterative process. As an alternative, we propose a simple yet novel method that leverages the latent space of the guidance network, which is designed as a variational autoencoder taking full-body poses as inputs. By enforcing latent distributions of this guidance network and our head-to-motion network to be similar, latent features sampled from the 'guided' distribution, i.e., distribution learned in our head-to-motion network, can be reliably decoded for natural representations of full-body poses even only with the head pose. One important advantage of the proposed method is that one-step sampling scheme achieves remarkably fast inference (more than 50 times faster) compared to diffusion-based approaches. Experimental results on benchmark datasets show that the proposed method efficiently improves the performance of ego-body mesh reconstruction.