发表机构
Sogang University(西江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出采样率无关的TF-Refiner模型,通过可配置分析带宽解耦深度处理与全频带,在多种采样率下超越特定率模型,支持推理时灵活选择成本-质量平衡。
AI 中文摘要
语音增强系统通常针对固定的采样率开发,而随着频率仓数量的增加,时频模型的成本也随之上升。我们提出了TF-Refiner,一种与采样频率无关的模型,它将深度分析带宽与全频带输入和输出解耦。一个深度编码器处理可配置截止频率以下的频带,而一个浅层解码器将编码特征与依赖于输入的高频带查询相结合,并预测应用于原始含噪STFT的局部复数滤波器。在16和48 kHz下训练的一组参数在各种采样率下进行评估。在VoiceBank+DEMAND数据集上,该通用模型在PESQ、STOI和对数谱距离方面,在评估的所有采样率(包括训练中未见过的采样率)上均优于特定于采样率的对应模型。随机截止训练使得在推理时无需重新训练或改变输出带宽即可选择成本-质量工作点。这些结果支持将可配置分析带宽作为多速率全频带增强的实用设计选择。
英文摘要
Speech enhancement systems are often developed for a fixed sampling rate, while time-frequency models become more expensive as the number of frequency bins increases. We propose TF-Refiner, a sampling-frequency-independent model that decouples the deep analysis bandwidth from the full-band input and output. A deep encoder processes the band below a configurable cutoff, while a shallow decoder combines the encoded features with input-dependent high-band queries and predicts local complex filters applied to the original noisy STFT. A single parameter set trained at 16 and 48 kHz is evaluated at various sampling rates. On VoiceBank+DEMAND, the universal model outperforms the rate-specific counterparts in PESQ, STOI, and log-spectral distance across the evaluated rates, including rates unseen in training. Random-cutoff training enables inference-time selection of cost-quality operating points without retraining or changing the output bandwidth. These results support configurable analysis bandwidth as a practical design choice for multi-rate full-band enhancement.
Comments5 pages, 5 figures, 4 tables. Submitted to ICASSP 2027