arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.06788cs.CVcs.SD

当语音遇见嘴唇:面向二语发音评估的可解释音视频同步

When Speech Meets Lips: Interpretable Audio-Visual Synchronization for L2 Pronunciation Assessment

Bowen Yu, Mingyu Huang, Yishen Liu, Yue Zhao

首次发表
浏览论文内容

中文总结 AI 辅助

针对二语发音评估忽视语音与发音动作时间同步的问题,提出可解释音视频同步框架,通过时滞轨迹和稳定性指数量化同步,连接自动评分与训练反馈。

中文摘要 AI 辅助

自动发音评估(APA)系统借助基于Transformer的模型和自监督语音表示已取得强劲性能。然而,大多数方法仅依赖声学信号,忽视了语音与发音动作之间的时间同步,从而限制了针对时间错配的诊断性反馈,而这对二语发音训练至关重要。我们提出了一种可解释的音视频同步框架,通过特征编码、交叉注意力融合、时滞估计、稳定性量化和可视化,显式建模语音与嘴唇的时间对齐。该框架引入了帧级时滞轨迹和时滞稳定性指数(LSI)来量化同步鲁棒性。我们还访谈了30名参与者,包括10名教师和20名具有不同母语背景的学生,以评估其有效性。通过将隐式对齐转化为可解释的表示,该框架将自动评分与可操作的计算机辅助发音训练反馈相连接。数据集和补充材料可在该https URL获取。

英文摘要

Automatic Pronunciation Assessment (APA) systems have achieved strong performance with transformer-based models and self-supervised speech representations. However, most methods rely only on acoustic signals and overlook temporal synchronization between speech and articulatory movements, limiting diagnostic feedback on timing mismatches important for L2 pronunciation training. We propose an interpretable audio-visual synchronization framework that explicitly models speech-lip temporal alignment through feature encoding, cross-attention fusion, lag estimation, stability quantification, and visualization. The framework introduces frame-level lag trajectories and a Lag Stability Index (LSI) to quantify synchronization robustness. We also interviewed 30 participants, including 10 instructors and 20 students with diverse first-language backgrounds, to assess its effectiveness. By transforming implicit alignment into interpretable representations, the framework connects automatic scoring with actionable Computer-Aided Pronunciation Training feedback. Datasets and supplemental materials are available at https://www.robots.ox.ac.uk/~vgg/data/lip_reading/.

发表机构

  • Wenzhou-Kean University(温州肯恩大学)

机构由 AI 辅助整理,请以论文原文为准。

↑