发表机构
Graduate Institute of Communication Engineering, National Taiwan University; Research Center for Information Technology Innovation, Academia Sinica; Department of Otolaryngology-Head and Neck Surgery, Taipei Veterans General Hospital; School of Medicine, National Yang Ming Chiao Tung University; Center for Hearing Research, University of California, Irvine(国立台湾大学通讯工程学研究所; 中央研究院资讯科技创新研究中心; 台北荣民总医院耳鼻喉头颈外科; 国立阳明交通大学医学院; 加州大学欧文分校听力研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对助听器在竞争语音下性能受限的问题,提出端到端视听框架AV-NeuroAMP,融合语音、视频和听力图,联合增强与个性化放大,实验证明优于传统方法。
AI 中文摘要
在噪声中理解语音对助听器用户而言仍然具有挑战性,尤其是在存在竞争说话者的情况下。传统助听器通常分阶段进行语音增强(SE)和听力损失补偿,这可能导致增强错误和信号失真传递到放大阶段。此外,在竞争语音条件下,仅音频的语音增强往往提供的益处有限,因为目标语音和干扰语音具有相似的声学特性,难以分离。为解决这些局限性,我们提出了AV-NeuroAMP,一种端到端的视听框架,它整合了噪声语音、目标说话者视频和听者的听力图,以联合执行语音增强、个性化放大和动态范围压缩。我们进一步引入了听力图条件的特征级线性调制(AC-FiLM),以有效纳入听者特定的听力特征。客观评估显示,AV-NeuroAMP在域内英语测试集上优于传统放大、仅音频的NeuroAMP模型以及两阶段系统,并且这些改进在未见过的普通话测试集上得以保持。涉及模拟听力损失下的正常听力参与者和有听力损失的参与者的听力测试进一步证明了语音质量和可懂度的改善,其中在竞争语音条件下观察到的益处最大。这些发现支持端到端视听个性化放大作为在具有挑战性的声学环境中提高助听器性能的一种有前景的方法。
英文摘要
Speech understanding in noise remains challenging for hearing-aid users, particularly in the presence of competing speakers. Conventional hearing aids typically perform speech enhancement (SE) and hearing-loss compensation in separate stages, which may cause enhancement errors and signal distortions to carry over to the amplification stage. Moreover, audio-only SE often provides limited benefits under competing-speech conditions because the target and interfering speech share similar acoustic characteristics, making them difficult to separate. To address these limitations, we propose AV-NeuroAMP, an end-to-end audiovisual framework that integrates noisy speech, target-talker video, and the listener's audiogram to jointly perform SE, personalized amplification, and dynamic-range compression. We further introduce audiogram-conditioned feature-wise linear modulation (AC-FiLM) to effectively incorporate listener-specific hearing profiles. Objective evaluations showed that AV-NeuroAMP outperformed conventional amplification, the audio-only NeuroAMP model, and two-stage systems on an in-domain English test set, and that these improvements were retained on an unseen Mandarin test set. Listening tests involving normal-hearing participants under simulated hearing loss and listeners with hearing loss further demonstrated improvements in speech quality and intelligibility, with the greatest benefits observed under competing-speech conditions. These findings support end-to-end audiovisual personalized amplification as a promising approach for improving hearing-aid performance in challenging acoustic environments.