发表机构
Department of Computer Science, University of Bari Aldo Moro; Institute for Technological Development and Innovation in Communications, University of Las Palmas de Gran Canaria(巴里阿尔多·莫罗大学计算机科学系; 拉斯帕尔马斯大加那利大学通信技术发展与创新研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究手语生成问题,提出物理信息扩散模型PIDiffSign,将解剖学约束融入架构与训练目标,用Transformer编码器 - 解码器及可微几何模块训练,实验证明该模型能提升手语生成的运动真实感和语义保真度。
AI 中文摘要
手语生成需同时满足语义保真度和生物力学合理性两个约束。现有方法通过基于坐标的目标优化语义重建,存在骨长漂移等问题。我们引入PIDiffSign,一种用于从词汇到姿势翻译的物理信息扩散模型,将解剖学约束纳入架构和训练目标。模型使用Transformer编码器 - 解码器,通过自适应零初始化层归一化等方式训练。实验表明该方法在多个方面有改进,证明其能提升手语生成的运动真实感和语义保真度。
英文摘要
Sign language production, which generates continuous 3D skeletal motion from spoken language input, must simultaneously satisfy two constraints: semantic fidelity, so that a deaf viewer can recognize the intended sequence of glosses, and biomechanical plausibility, so that the generated skeleton respects anatomical constraints. Existing approaches optimize semantic reconstruction through coordinate-based objectives that treat the skeleton as an unstructured vector, thus allowing for bone length drift, joint angle violations, and temporarily locked fingers. We introduce PIDiffSign, a physics-informed diffusion model for gloss-to-pose translation that incorporates anatomical constraints into both the architecture and training objective. The model uses a Transformer encoder-decoder, where the decoder is conditioned on the diffusion time step through adaptive zero-initialized layer normalization and cross-attends to gloss representations. A differentiable geometry module enforces bone length consistency and biologically valid joint angles throughout generation. Training combines anthropomorphic, kinematic, angular, and finger-joint constraints with a contrastive gloss-pose alignment loss and classifier-free guidance for semantically conditioned sampling. Experiments on the PHOENIX14T and CSL-Daily benchmarks show consistent improvements over a strong diffusion baseline in pose accuracy, joint-angle correctness, distributional realism, and back-translation quality. These results demonstrate that physics-informed diffusion improves both motion realism and semantic fidelity for sign language generation.