arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29443cs.CV

姿态自适应动态FiLM调制用于视觉语音识别

Pose Adaptive Dynamic FiLM Modulation for Visual Speech Recognition

Matthew Kit Khinn Teng, Haibo Zhang, Takeshi Saitoh

首次发表
浏览论文内容

中文总结 AI 辅助

针对视觉语音识别中头部姿态变化问题,提出姿态自适应动态FiLM框架,通过DR-FiLM预测输入依赖权重动态控制调制强度,在LRS2和LRS3上显著降低音素错误率。

中文摘要 AI 辅助

头部姿态变化在视觉语音识别(VSR)中引入了显著的外观变换,因此姿态感知的特征调制是可取的。然而,使用大量具有固定调制强度的特征级线性调制(FiLM)电路可能导致性能下降和不期望的特征交互。我们提出了一种姿态自适应动态FiLM框架,其中包含动态残差FiLM(DR-FiLM)调制器,该调制器预测依赖于输入的权重,以自适应地控制姿态条件调制的强度。在LRS2和LRS3上的实验表明,未加权的多路径调制显著降低了音素识别性能,与单一ResFiLM配置的16.20%和20.96%相比,PER分别增加到20.33%和29.42%。相比之下,所提出的带有动态Deep-Res加权的DR-FiLM在LRS2上将PER降低到15.74%,在LRS3上降低到23.91%,显著减轻了未加权调制的不利影响。对学习权重的分析进一步揭示了一种一致的倾向:随着头部姿态变化的增加,将更大的权重分配给更深的FiLM路径。这些结果表明,当调制强度被动态控制时,合并姿态条件FiLM电路更为有效。

英文摘要

Head-pose variation introduces substantial appearance transformations in visual speech recognition (VSR), making pose-aware feature modulation desirable. However, performance degradation and unwanted feature interactions may result from using numerous Feature-wise Linear Modulation (FiLM) circuits with fixed modulation intensity. We propose a Pose Adaptive Dynamic FiLM framework with a Dynamic Residual FiLM (DR-FiLM) modulator that predicts input-dependent weights to adaptively control the strength of pose-conditioned modulation. Experiments on LRS2 and LRS3 demonstrate that unweighted multi-pathway modulation substantially degrades phoneme recognition, increasing PER to 20.33% and 29.42%, respectively, compared with 16.20% and 20.96% for the single ResFiLM configuration. In contrast, the proposed DR-FiLM with dynamic Deep-Res weighting reduces PER to 15.74% on LRS2 and 23.91% on LRS3, substantially mitigating the adverse effects of unweighted modulation. The analysis of the learned weights further reveals a consistent tendency to assign greater weight to the deeper FiLM pathway as head-pose variation increases. These results show that merging pose-conditioned FiLM circuits is more efficient when the modulation strength is dynamically controlled.

发表机构

  • Kyushu Institute of Technology(九州工业大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑