arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SEER:面向人工耳蜗语音的基于检索的源条件情感增强

SEER: Source-Conditioned Emotion Enhancement via Retrieval for Cochlear-Implant Speech

Hsing-Hang Chou, Yun-Shao Lin, Ching-Chin Sung, Chi-Chun Lee

arXiv 2610.11131首次发表:更新:

AI 中文总结

针对人工耳蜗语音情感识别线索被削弱的问题,提出无需并行录音或强度标签的SEER检索框架,在RAVDESS、ESD数据集及多条件下显著提升情感识别性能。

AI 中文摘要

人工耳蜗(CI)可恢复言语感知,但会削弱语音情感识别所需的线索。现有面向CI的增强方法需要并行的正常/强录音及强度标签。我们提出SEER,一种基于检索的框架,用于学习CI处理后哪种相同情感的参考能帮助每个源保持可识别性。源条件检索器从采样的情感语音转换结果中学习CI感知效用,同时不确定性引导探索避免穷尽配对评估;无需并行录音或强度标签。SEER在RAVDESS数据集上将N8下的源宏F1提升7.30个百分点,在ESD数据集上提升11.66个百分点,且ESD在N4/N8/N16下均有显著提升;16名听众的RAVDESS在所有条件下均有显著提升。穷尽分析发现,更强参考带来总体收益,但性别或内容匹配几乎无影响。

英文摘要

Cochlear implants (CIs) restore speech access but weaken cues needed for vocal emotion recognition. Prior CI-oriented enhancement requires parallel normal/strong recordings and intensity labels. We propose SEER, a retrieval-based framework that learns which same-emotion reference helps each source remain recognizable after CI processing. A source-conditioned retriever learns CI-aware utility from sampled emotional voice conversion outcomes, while uncertainty-guided exploration avoids exhaustive pair evaluation; neither parallel recordings nor intensity labels are required. SEER improves Source macro-F1 at N8 by 7.30 points on RAVDESS and 11.66 points on ESD, with significant ESD gains across N4/N8/N16. Sixteen-listener RAVDESS gains are significant across all conditions. Exhaustive analysis finds an aggregate benefit from stronger references but little effect from matching gender or content.

CommentsSubmitted to ICASSP 2027. 5 pages, 2 figures, 3 tables. Code and demo audio: https://github.com/HenryChou36/SEER

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑