arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.23846cs.SDeess.AS

基于语义嵌入的音频自动均衡

Automatic Audio Equalization with Semantic Embeddings

Eloi Moliner, Vesa Välimäki, Konstantinos Drossos, Matti S. Hämäläinen

首次发表
浏览论文内容

中文总结 AI 辅助

研究提出基于语义嵌入的数据驱动方法实现音频自动盲均衡,用预训练模型提供语义嵌入,训练轻量级头部,在音乐和语音上训练,经评估验证其有效性,凸显在现实音频增强应用中的潜力。

中文摘要 AI 辅助

本文提出一种数据驱动方法,通过预测对数梅尔频谱特征并推导逆滤波器来实现音频的自动盲均衡。该方法使用深度神经网络,以预训练模型提供的语义嵌入为骨干,仅训练一个轻量级头部,旨在提高训练效率和泛化能力。模型在音乐和语音上训练,对噪声和混响具有鲁棒性。客观评估证实其有效性,主观测试表明性能与使用真实对数梅尔频谱特征的神谕相当,虽滤波阶段有局限,但结果凸显该方法在现实音频增强应用中的潜力。

英文摘要

This paper presents a data-driven approach to automatic blind equalization of audio by predicting log-mel spectral features and deriving an inverse filter. The method uses a deep neural network, where a pre-trained model provides semantic embeddings as a backbone, and only a lightweight head is trained. This design is intended to enhance training efficiency and generalization. Trained on both music and speech, the model is robust to noise and reverberation. Objective evaluations confirm its effectiveness, and subjective tests show performance comparable to that of an oracle that uses true log-mel spectral features, indicating that the model accurately estimates the desired characteristics, with remaining limitations attributed to the filtering stage. Overall, the results highlight the potential of the method for real-world audio enhancement applications.

发表机构

  • Aalto University(阿尔托大学)
  • Nokia Technologies(诺基亚技术公司)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑