AI 中文总结
该研究针对视听语音识别跨模态融合缺乏结构化迭代优化的问题,提出DoubleHelix融合框架,通过三个组件实现迭代交互与退化感知增强,在LRS3数据集上大幅降低词错误率并提升噪声鲁棒性。
AI 中文摘要
视听语音识别(AVSR)依赖于音频与视觉模态的有效融合,但现有方法将跨模态交互视为单步操作,缺乏结构化迭代优化。本文提出DoubleHelix多模态融合框架,将融合重构为带自适应退化感知增强的迭代跨模态交互过程,包含三个组件:ReverseParallelHelix用于带学习对齐约束的多轮结构化交互,QualitySensor用于学习退化感知门控信号,HelixReplication用于一致性引导的条件特征增强。在LRS3数据集上的实验显示,DoubleHelix在纯净音频下的词错误率(WER)为0.68%,在匹配主干网络设置下较此前最优结果实现5.6%的相对提升;全面的 ablation 研究验证了各组件的贡献,包括非对称通路权重等设计选择的针对性分析。该框架在评估的 babble 噪声条件下展现出更强的鲁棒性,在信噪比(SNR)为-5dB时达到11.6%的WER。
英文摘要
Audio-visual speech recognition (AVSR) relies on effective fusion of audio and visual modalities, yet existing approaches treat cross-modal interaction as a single-step operation without structured iterative refinement. We present DoubleHelix, a multimodal fusion framework that reformulates fusion as an iterative cross-modal interaction process with adaptive degradation-aware enhancement. The framework comprises three components including ReverseParallelHelix for multi-turn structured interaction with learned alignment constraints, QualitySensor for learning degradation-aware gating signals, and HelixReplication for consistency-guided conditional feature enhancement. Experiments on LRS3 demonstrate that DoubleHelix achieves 0.68% WER on clean audio, outperforming previous best results by 5.6% relative improvement under matched backbone settings. Comprehensive ablation studies validate each component contribution, including targeted analysis of design choices such as asymmetric pathway weighting. The framework shows improved robustness under evaluated babble-noise conditions, achieving 11.6% WER at SNR -5dB.
CommentsACM MM2026 ACCEPTED