arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

KoUniTalk:一个轻量级以发音为中心的韩英3D说话人脸基准

KoUniTalk: A Lightweight Articulation-Centered Korean-English 3D Talking Face Benchmark

Hyunjung Chung, Unsang Park

arXiv 2609.19840首次发表:更新:

发表机构

Sogang University(西江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

KoUniTalk提出一个轻量级韩英3D说话人脸基准,通过形变迁移统一网格拓扑,将维度降至3,528,支持跨语言语音驱动面部动画训练与评估。

AI 中文摘要

高质量的3D说话人脸数据集在很大程度上仍以英语为中心,而韩语3D面部运动数据由于网格拓扑、空间尺度、坐标系和时间采样等方面的差异,难以与标准英语基准相结合。我们提出了KoUniTalk,一个轻量级以发音为中心的韩英3D说话人脸基准,它通过形变迁移将VOCASET和已发布的基于韩语语音的3D说话人脸数据重定向到共享的网格拓扑。KoUniTalk并非提出新的形变迁移算法或全头身份保持的虚拟形象数据集,而是提供了一个身份中立的规范输出空间,用于在英语和韩语之间进行受控的语音驱动面部发音训练和评估。统一模板包含1,176个顶点,并聚焦于嘴部及相邻的下脸和中脸区域,将输出维度从15,069和72,147维降至3,528维,与VOCASET/FLAME和原始韩语网格相比,分别减少了4.27倍和20.45倍。为了检验重定向是否保留了与语音相关的运动,我们评估了语义嘴部地标轨迹,包括嘴部开合、嘴部宽度、开合比和嘴部开合动态。由于韩语数据集的官方测试集未公开,我们额外定义了一个主体不相交的韩语基准划分。处理后的匹配基准包含22个说话人、4,978个序列和642,781帧,使得在单一紧凑的发音模板空间中,能够进行韩英跨数据集的语音驱动3D面部动画模型评估。来源报告的库存计数与这些处理后的计数分开列出。

英文摘要

High-quality 3D talking face datasets remain largely English-centric, and Korean 3D facial motion data are difficult to combine with standard English benchmarks because of differences in mesh topology, spatial scale, coordinate system, and temporal sampling. We present KoUniTalk, a lightweight articulation-centered Korean-English 3D talking face benchmark that retargets VOCASET and the released Korean speech-based 3D talking face data to a shared mesh topology using deformation transfer. Rather than proposing a new deformation-transfer algorithm or a full-head identity-preserving avatar dataset, KoUniTalk provides an identity-neutral canonical output space for controlled speech-driven facial articulation training and evaluation across English and Korean. The unified template contains 1,176 vertices and focuses on the mouth and adjacent lower- and mid-face regions, reducing the output dimensionality from 15,069 and 72,147 dimensions to 3,528 dimensions, corresponding to 4.27-fold and 20.45-fold reductions compared with VOCASET/FLAME and the original Korean mesh, respectively. To examine whether retargeting preserves speech-relevant motion, we evaluate semantic mouth-landmark trajectories, including mouth opening, mouth width, aperture ratio, and mouth-opening dynamics. Since the official test set of the Korean dataset is not publicly released, we additionally define a subject-disjoint Korean benchmark split. The processed matched benchmark contains 22 speakers, 4,978 sequences, and 642,781 frames, enabling Korean-English cross-dataset evaluation of speech-driven 3D facial animation models in a single compact articulation-template space. Source-reported inventory counts are listed separately from these processed counts.

Comments22 pages, 5 figures; includes supplementary material

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑