arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.08496cs.CV

SignRefine:适配基础视频模型用于手语生成

SignRefine: Adapting Foundational Video Models for Sign Language Generation

  • University of Surrey(萨里大学)
  • Centre for Vision, Speech and Signal Processing(视觉、语音与信号处理中心)

机构由 AI 辅助整理,请以论文原文为准。

Anton Pelykh, Edward Fish, Ozge Mercanoglu Sincan, Richard Bowden

AI总结:

SignRefine通过局部适配器优化手部面部动作,基于关键点生成可理解手语视频,配合NVSign数据集,手部精度提升30%,用户偏好超80%。

AI中文摘要:

手语视频生成要求精确的手部和面部动作,然而现代视频扩散模型主要在口语视频上训练,会产生使手语难以理解的人工痕迹。我们提出SignRefine,一种仅通过2D关键点条件即可生成可理解手语的手语视频生成模型,能够跨外观和视觉条件进行泛化。我们的方法基于预训练的视频扩散Transformer,并引入具有空间定位的局部适配器,选择性地优化手部和面部区域,将强大基础模型的先验引导至精确的动作生成。为支持此项工作并促进更广泛的手语研究,我们提出了NVSign,一个大规模的手语原生视频内容数据集,提供多样的手语者外观、环境和自然对话场景。在该数据上训练后,我们的模型在手部姿态精度指标上相比最强基线提升了高达30%,并且超过80%的比较中,手语使用者更偏好我们模型的视觉质量和可理解性。

英文摘要:

Sign language video generation demands precise hand and facial articulation, yet modern video diffusion models, trained predominantly on spoken-language video, produce artifacts that render signing unintelligible. We propose SignRefine, a sign language video generation model that produces comprehensible signing from 2D keypoint conditioning alone, generalizing across appearances and visual conditions. Our approach builds on a pretrained video diffusion transformer and introduces local adapters with spatial grounding to selectively refine hand and face regions, steering the strong base model's prior toward accurate articulation. To enable this work and support broader sign language research, we present NVSign, a large-scale dataset of video content natively produced in sign language, offering diverse signer appearances, environments, and natural conversational settings. Trained on this data, our model shows up to 30% improvement in hand pose precision metrics over the strongest baseline and is preferred by sign language users for visual quality and comprehensibility in more than 80% of comparisons.

↑