arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

文本到手语:用于文本到手语视频生成的单 GPU 扩散基线

Text2Sign: A Single-GPU Diffusion Baseline for Text-to-Sign Language Video Generation

Ruize Xia

arXiv 2607.13164首次发表:更新:

AI 中文总结

研究文本到手语视频生成,提出Text2Sign模型,结合冻结视觉语言文本编码器等技术,降低全视频注意力成本。在特定分割上有一定验证损失等结果,虽提示敏感性弱,但为单GPU研究提供了基线。

AI 中文摘要

手语是数百万聋人和听力障碍者的主要交流渠道,但文本到手语视频生成成本高昂,因为视频扩散模型训练和评估成本高。本文提出Text2Sign,一种用于短手语片段的文本条件扩散模型,可在单个NVIDIA L4 GPU上运行。它将冻结的视觉语言文本编码器与3D编码器 - 解码器和分解的时空注意力相结合,以降低全视频注意力成本并保持运动连贯性。我们比较了仅卷积和Transformer风格的主干、冻结的预训练和特定任务的文本编码器以及分解与全注意力。在签名者不相交的How2Sign分割上,最佳短期消融达到验证损失0.0648,而长期检查点达到0.00999。在紧凑评估切片上,使用8步DDIM采样和5.0的引导尺度,后者实现了0.2403±0.0238的SSIM、15.11±0.42 dB的PSNR和1.0000±0.0000的时间一致性。它在12.60秒内生成32帧、64×64的片段,即每秒2.54帧,峰值推理内存为3.12 GB。保留的去噪审计显示提示敏感性较弱:删除文本会使后期时间步损失从0.9875增加到0.9891,而打乱的提示与正确提示表现相似。因此,冻结的文本条件改善了短预算验证损失,但提示特定的分离仍然有限。该系统限于低分辨率、短片,缺乏专家语言评估,应视为单GPU研究基线而非完整的手语生产系统。代码可在该https URL获取。

英文摘要

Sign language is a primary communication channel for millions of Deaf and hard-of-hearing people, yet text-to-signer video generation remains costly because video diffusion models are expensive to train and evaluate. This paper presents Text2Sign, a text-conditioned diffusion model for short sign-language clips that runs on a single NVIDIA L4 GPU. It combines a frozen vision-language text encoder with a 3D encoder-decoder and factorized spatiotemporal attention to reduce the cost of full-video attention while preserving motion coherence. We compare convolution-only and transformer-style backbones, frozen pretrained and task-specific text encoders, and factorized versus full attention. On a signer-disjoint How2Sign split, the best short-run ablation reaches a validation loss of 0.0648, while a longer-run checkpoint reaches 0.00999. On a compact evaluation slice, the latter achieves an SSIM of $0.2403 \pm 0.0238$, a PSNR of $15.11 \pm 0.42$ dB, and temporal consistency of $1.0000 \pm 0.0000$ using 8-step DDIM sampling with a guidance scale of 5.0. It generates a 32-frame, $64 \times 64$ clip in 12.60 seconds, or 2.54 frames per second, with peak inference memory of 3.12 GB. A held-out denoising audit shows only weak prompt sensitivity: removing text increases late-timestep loss from 0.9875 to 0.9891, while shuffled prompts perform similarly to correct prompts. Frozen text conditioning therefore improves short-budget validation loss, but prompt-specific separation remains limited. The system is restricted to low-resolution, short clips and lacks expert linguistic evaluation, so it should be viewed as a single-GPU research baseline rather than a complete sign-language production system. Code is available at https://github.com/xiaruize0911/text2sign.

Journal refIEEE Access, vol. 14, pp. 64003-64017, 2026

DOI:10.1109/ACCESS.2026.3686260

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑