SignRR:检索并优化真实手语动作以实现手语生成
SignRR: Retrieve and Refine Real Motion for Sign Language Production
浏览论文内容
中文总结 AI 辅助
针对手语生成的现有范式局限,提出 SignRR 框架,采用检索并优化的范式,在 PHOENIX14T 和 CSL-Daily 上实现了最优回译性能且姿态质量具竞争力。
中文摘要 AI 辅助
手语生成(Sign Language Production,SLP)旨在从口语生成连续的手语动作,通常通过 gloss-to-pose 生成实现。现有工作主要遵循两种范式:生成模型从学习到的先验或噪声中合成动作,无需参考观察到的手语实例,难以保留罕见的手部构型和手语者特定的发音;基于检索的方法复用真实、发音良好的动作片段,但拼接不同手语者和协同发音语境的片段会引入整个序列的节奏和风格不一致,不仅出现在片段边界。这些局限表明互补解决方案:用检索提供真实发音,用学习到的优化施加仅靠检索缺失的全局一致性。因此提出检索并优化范式,从真实检索到的动作开始,将其优化为全局一致的手语序列,而非从头生成动作。本框架 SignRR 从真实手语片段字典初始化动作,通过部分感知残差 VQ-VAE 优化整个序列,其中残差量化保留精细手部发音,时间长度差异在潜在空间处理。在 PHOENIX14T 和 CSL-Daily 上的实验表明,SignRR 实现了最先进的回译性能,同时保持有竞争力的姿态质量。
英文摘要
Sign language production (SLP) aims to generate continuous signing motion from spoken language, often through gloss-to-pose generation. Prior work mainly follows two paradigms. Generative models synthesize motion from a learned prior or from noise, without reference to an observed signing instance, making rare hand configurations and signer-specific articulation difficult to preserve. Retrieval-based methods reuse real, well-articulated motion segments, but concatenating segments from different signers and co-articulation contexts can introduce rhythm and style inconsistencies across the full sequence, not only at segment boundaries. These limitations suggest a complementary solution: use retrieval to provide realistic articulation, and use learned refinement to impose the global coherence that retrieval alone lacks. We therefore propose retrieve-and-refine, a paradigm that starts from real retrieved motion and refines it into a globally coherent signing sequence rather than generating motion from scratch. Our framework, SignRR, initializes motion from a dictionary of real sign segments and refines the full sequence with a part-aware Residual VQ-VAE, where residual quantization preserves fine hand articulation and temporal length differences are handled in the latent space. Experiments on PHOENIX14T and CSL-Daily show that SignRR achieves state-of-the-art back-translation performance while maintaining competitive pose quality.
发表机构
- University of Central Florida(中佛罗里达大学)
- Universidad Nacional Mayor de San Marcos(圣马科斯国立大学)
- Universidad Catolica San Pablo(圣巴勃罗天主教大学)
- Marist University(玛瑞斯特大学)
机构由 AI 辅助整理,请以论文原文为准。