arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FlexiAvatar:任意身体可见性下的统一3D高斯人体头像

FlexiAvatar: Unified 3D Gaussian Human Avatars Under Arbitrary Body Visibility

Yihalem Yimolal Tiruneh, Muhammad Salman Ali, Uyoung Jeong, Muneeb A. Khan, MD Khalequzzaman Chowdhury Sayem, Allanur Bayramgeldiyev, Binod Bhattarai, Seungryul Baek

arXiv 2607.19100首次发表:更新:

发表机构

UNIST; University of Aberdeen; University College London; Fogsphere(韩国蔚山科学技术院; 阿伯丁大学; 伦敦大学学院; 雾球公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究从单目视频重建3D人体头像问题,提出FlexiAvatar框架,结合遮挡鲁棒跟踪与特定部分残差细化捕捉细节,用扩散方法处理不可见区域,实验表明其重建质量高,还能减少运行时和内存开销。

AI 中文摘要

从单目视频重建可动画的3D人体头像是计算机视觉中的一个基本问题,在AR/VR和数字内容创作中有广泛应用。现有方法通常将参数化身体模型与神经渲染或3D高斯点云拼接相结合,从短视频中联合优化所有身体区域,这往往会降低可见区域的保真度。为克服这一限制,我们引入了FlexiAvatar,这是一个统一框架,仅显式优化可见身体区域,有效消除未观察到的肢体产生的伪影。我们的方法将遮挡鲁棒的SMPL-X跟踪与特定部分的残差细化相结合,以捕捉高频几何和外观细节。为完成完全不可见的区域(如后视图),我们利用基于扩散的方法生成与观察到的外观一致的纹理。在全身(NeuMan、ZJU-MoCap、WildAvatar)、上半身/半身(脱口秀片段)和仅头部(INSTA)输入上的实验表明,FlexiAvatar始终提供更高的重建质量,在各数据集上平均PSNR提高约3%,优于现有方法。最后,通过将优化限制在观察到的区域,我们的方法减少了必须优化和渲染的高斯点的有效数量,从而在部分可见性场景中减少了运行时和内存开销。

英文摘要

Reconstructing animatable 3D human avatars from monocular video is a fundamental problem in computer vision with broad applications in AR/VR and digital content creation. Existing approaches typically couple parametric body models with neural rendering or 3D Gaussian splatting and optimize all body regions jointly from short videos, which often degrades fidelity in the visible areas. To overcome this limitation, we introduce FlexiAvatar, a unified framework that explicitly optimizes only the visible body regions, effectively eliminating artifacts arising from unobserved limbs. Our method integrates occlusion-robust SMPL-X tracking with part-specific residual refinement to capture high-frequency geometric and appearance details. To complete entirely unseen regions (e.g., back views), we leverage a diffusion-based approach to generate texture consistent with the observed appearance. Experiments on full-body (NeuMan, ZJU-MoCap, WildAvatar), upper/half-body (talk-show clips), and head-only (INSTA) inputs show that FlexiAvatar delivers consistently higher reconstruction quality, outperforming state-of-the-art methods by an average PSNR improvement of approximately 3% across datasets. Finally, by restricting optimization to observed regions, our method reduces the effective number of Gaussians that must be optimized and rendered, leading to reduced runtime and memory overhead in partial-visibility scenarios.

CommentsAccepted in ECCV 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑