SeRV:面向美国手语生成的语义对齐残差向量量化
SeRV: Semantic-Aligned Residual Vector Quantization for American Sign Language Generation
- University of Oklahoma(俄克拉荷马大学)
- University of Georgia(佐治亚大学)
- University of Massachusetts Amherst(马萨诸塞大学阿默斯特分校)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
针对ASL生成中语义一致性不足的问题,提出语义对齐残差向量量化分词器SeRV,结合层次化GPT从文本生成3D动作,在375小时视频上取得最先进姿态精度。
中文摘要 AI 辅助
美国手语(ASL)生成仍具挑战性,原因在于配对的文本-ASL动作数据有限,且难以学习既精确用于重建又能从语言输入中预测的动作表示。现有方法依赖为重建优化的动作分词器,缺乏来自配对文本的显式语义监督。因此,学习到的分词在支持语义一致且细粒度的ASL动作生成方面仍有限。为解决此局限,我们提出SeRV(语义对齐残差向量量化),一种用于ASL生成的语义对齐RVQ分词器。SeRV通过结合句子级动作-文本对齐与词元级文本条件监督,学习语义结构化的残差词元空间。基于该分词器,层次化GPT以从粗到细的方式预测残差动作词元,生成结构连贯且语义对齐的3D ASL动作。我们进一步通过从YouTube-ASL视频中恢复配对的3D动作,构建大规模重建的3D ASL动作-文本基准。在375小时ASL视频上的实验表明,SeRV在How2Sign和YouTube-ASL数据集上均达到最先进的姿态精度,同时直接从文本生成语义一致的3D ASL动作。
英文摘要
American Sign Language (ASL) generation remains challenging due to limited paired text-ASL motion data and the difficulty of learning motion representations both precise for reconstruction and predictable from linguistic input. Existing methods rely on motion tokenizers optimized for reconstruction, without explicit semantic supervision from paired text. As a result, the learned tokens remain limited in supporting semantically consistent and fine-grained ASL motion generation. To address this limitation, we propose SeRV (Semantic-Aligned Residual Vector Quantization), a semantic-aligned RVQ tokenizer for ASL generation. SeRV learns a semantically structured residual token space by combining sentence-level motion-text alignment with token-level text-conditioned supervision. Building on this tokenizer, a Hierarchical GPT predicts residual motion tokens in a coarse-to-fine manner, generating structurally coherent and semantically aligned 3D ASL motion. We further construct a large-scale reconstructed 3D ASL motion-text benchmark by recovering paired 3D motion from YouTube-ASL videos. Experiments across 375 hours of ASL video show that SeRV achieves state-of-the-art pose accuracy on both How2Sign and YouTube-ASL datasets, while producing semantically consistent 3D ASL motion directly from text.