arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OASIS:基于3D高斯溅射的遮挡感知单图像手部化身重建

OASIS: Occlusion-aware Single-image Hand Avatar Reconstruction via 3D Gaussian Splatting

Zhisheng Han, Shiyao Wu, Jiayan Qiu, Yakun Ju, Lu Liu, Le Zhang, Pengfei Feng, Huiyu Zhou, Zheheng Jiang

arXiv 2607.29633首次发表:更新:

发表机构

University of Leicester; University of Exeter; University of Birmingham; China University of Geoscience(莱斯特大学; 埃克塞特大学; 伯明翰大学; 中国地质大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出OASIS框架,通过3D高斯溅射结合可见性条件注意力与网格上特征表示,实现单图像手部化身重建,在视觉保真度、效率及下游应用通用性上优于现有方法。

AI 中文摘要

单图像3D手部化身重建本质上是不适定问题,且因高度铰接的手部存在严重自遮挡导致视觉证据有限,以及依赖姿态的复杂变形而极具挑战性。现有方法主要依赖隐式NeRF风格表示,其体积拟合计算成本高,且常难以保留手部细粒度细节。本研究提出OASIS,一种用于单图像手部化身重建的定制化3D高斯溅射框架。为在单视图重建中忠实地编码稀疏的图像特定外观线索,我们通过将输入图像观测与3D手部几何明确对齐,并对所得视觉证据进行上下文自适应分词,构建几何对齐的视觉证据令牌。由于严重自遮挡使图像证据的可靠性固有地依赖可见性,我们引入可见性条件的点-图像注意力,以可靠地将视觉证据传递给几何令牌,从而生成用于忠实且鲁棒重建的遮挡感知高斯特征。为进一步捕捉铰接手部的非刚性变形,我们引入网格上特征表示,使高斯变形能由局部表面拉伸引导。在该框架下,我们采用一次性适应方案,从多身份训练数据中学习共享手部先验,随后将其适配到目标图像以实现目标特定重建。大量实验表明,在挑战性姿态和野外场景中,OASIS在视觉保真度和效率上均优于现有基线,且在文本到化身生成、纹理编辑等下游应用中展现出强大通用性。

英文摘要

Single-image 3D hand avatar reconstruction is fundamentally ill-posed and particularly challenging due to limited visual evidence under severe self-occlusion and the complex pose-dependent deformation of highly articulated hands. Existing methods predominantly rely on implicit NeRF-style representations, whose volumetric fitting is computationally expensive and often struggles to preserve fine-grained hand details. In this work, we present OASIS, a tailored 3D Gaussian Splatting framework for single-image hand avatar reconstruction. To faithfully encode sparse image-specific appearance cues in single-view reconstruction, we construct geometry-aligned visual evidence tokens by explicitly aligning input image observations with 3D hand geometry and context-adaptively tokenizing the resulting visual evidence. Since severe self-occlusion makes the reliability of image evidence inherently visibility-dependent, we introduce a visibility-conditioned point-image attention to reliably transfer visual evidence to geometric tokens, yielding occlusion-aware Gaussian features for faithful and robust reconstruction. To further capture non-rigid deformation of articulated hands, we introduce a Feature-on-Mesh representation to enable Gaussian deformation to be guided by local surface stretching. Under this framework, we adopt a one-shot adaptation scheme that learns a shared hand prior from multi-identity training data and then fits it to a target image for target-specific reconstruction. Extensive experiments show that OASIS outperforms existing baselines in both visual fidelity and efficiency across challenging poses and in-the-wild scenarios, and further demonstrates strong versatility in downstream applications such as text-to-avatar generation and texture editing.

CommentsAccepted to ACM Multimedia 2026. Project page: https://mova-hand.github.io/MOVA/. Code repository: https://github.com/ivyyy77/OASIS

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑