arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过联合嵌入预测架构对齐实现用于语音同步的同步三维声道运动

Synchronized Three-Dimensional Vocal-Tract Motion for Speech Synchronization via Joint-Embedding Predictive Architecture Alignment

Sheng Li, Takahiro Shinozaki

arXiv 2607.11772首次发表:更新:

AI 中文总结

研究如何实现语音同步的三维声道运动,采用联合嵌入预测架构和强化学习/交叉熵方法,结合时长保持声学载体与校正的声道模型,在包含24个最小对刺激的评估中取得良好的自动语音识别等结果。

AI 中文摘要

现代神经语音系统能生成可理解的波形,但隐藏了产生声音的物理语音生成状态。生物力学声道模型虽能展现发音结构等,但直接物理波形合成不如现代神经声码器稳健。本文采用时长保持声学载体提供听觉波形,校正后的三维声道模型提供同步的颌、唇、舌等运动。通过联合嵌入预测架构(JEPA)风格表示和强化学习/交叉熵方法(RL/CEM)轨迹选择循环,使发音动作与声学载体及物理合理性约束对齐。评估包含12个3D录音,涵盖24个最小对刺激。在24词集上,载体取得了良好的自动语音识别(ASR)结果等多项指标。

英文摘要

Modern neural speech systems can generate intelligible waveforms, but they usually hide the physical speech-production state that produced the sound. Conversely, biomechanical vocal-tract models expose articulatory structure, contact behavior, airflow routing, and geometric constraints, but direct physical waveform synthesis remains less robust than modern neural vocoders. A duration-preserving acoustic carrier supplies the listening waveform, while a corrected three-dimensional vocal-tract model supplies synchronized jaw, lip, tongue, velum, laryngeal, oral-airflow, and nasal-airflow motion. A joint-embedding predictive architecture (JEPA)-style representation and a reinforcement learning/cross-entropy method (RL/CEM) trajectory-selection loop align articulatory actions to the acoustic carrier and to physical-plausibility constraints. The evaluation contains 12 3D recordings covering 24 minimal-pair stimuli. On the 24-word set, the carrier obtains good automatic speech recognition (ASR) results (an 8.33\% WER, a 4.17\% CER), a UTMOS score of 3.174, a mean JEPA score of 0.864, and a mean timbre-guard score of 0.947.

Commentspaper submitted to IEEE-SLT2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑