arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PHOSA:逼真三维手语虚拟形象建模与基准

PHOSA: Photorealistic 3D Sign Avatar Modeling and Benchmark

Haodong Wang, Hezhen Hu, Wengang Zhou, Houqiang Li

arXiv 2609.29292首次发表:更新:

发表机构

University of Science and Technology of China; University of Texas at Austin(中国科学技术大学; 德克萨斯大学奥斯汀分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出PHOSA方法,通过构建多视角中文手语数据集MVSign和解耦虚拟形象表示,实现高保真手语建模,并在手部与面部细节上取得优异效果,且能泛化至单目视频。

AI 中文摘要

在这项工作中,我们聚焦于逼真手语虚拟形象建模,这对于与聋人社区的有效沟通至关重要,其特点是复杂的手势和细微的面部表情。为此,我们引入了MVSign,这是首个与聋人专家共同设计的多视角中文手语数据集,具有多样化的手势和丰富的标注。为了实现精确的SMPL-X标注,我们开发了一种混合拟合流程,能够生成准确的躯体、手部和面部参数,并且也可应用于单目设置。基于MVSign,我们提出了一种解耦的手语虚拟形象表示,将躯体、头部和手部组件分离以捕捉复杂的关节运动,同时采用运动感知采样策略来处理运动模糊并平衡手势多样性。大量实验表明,我们的方法在MVSign上实现了高保真视觉效果,特别是在手部和面部细节区域,并且能够很好地泛化到野外单目手语视频。项目页面:此HTTPS URL。

英文摘要

In this work, we focus on photorealistic sign avatar modeling, which is crucial for effective communication with the Deaf community and is characterized by complex hand gestures and nuanced facial expressions. To this end, we introduce MVSign, the first multi-view Chinese sign language dataset co-designed with Deaf experts, featuring diverse gestures and rich annotations. For precise SMPL-X annotation, we develop a hybrid fitting pipeline that produces accurate body, hand, and facial parameters and can also be applied to the monocular setting. Building on MVSign, we propose a decoupled sign avatar representation that isolates body, head, and hand components to capture complex articulations, together with a motion-aware sampling strategy to handle motion blur and balance gesture diversity. Extensive experiments demonstrate that our method achieves high-fidelity visual results on MVSign, particularly in detailed hand and facial regions, and generalizes well to in-the-wild monocular sign language videos. Project page: https://naaapi.github.io/PHOSA.

CommentsECCV 2026, project page: https://naaapi.github.io/PHOSA

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑