arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

少样本校准用于仿真到真实单声道说话人距离估计

Few-Shot Calibration for Sim-to-Real Single-Channel Speaker Distance Estimation

Michael Neri, Archontis Politis, Tuomas Virtanen

arXiv 2609.29203首次发表:更新:

发表机构

Tampere University(坦佩雷大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对模拟训练说话人距离估计器在真实场景迁移差的问题,提出用少量真实标注数据通过事后校准映射调整输出,并基于排序能力选择合成检查点,实现有效校准。

AI 中文摘要

说话人距离估计器几乎完全在模拟房间声学上训练,因为带有真实说话人至麦克风距离标注的真实录音十分稀缺。我们表明,以这种方式训练的模型迁移效果不佳。在我们评估的三个真实语料库上,简单地预测语料库的平均距离比任何学习模型都更准确。然后,我们探究需要多少带标签的真实话语才能使一个冻结的、合成训练的估计器变得有用,并研究事后校准映射,这些映射在无需梯度或重新训练的情况下重新缩放其输出。对可实现误差的分析表明,校准不受限于估计器的绝对精度,而是受限于其按距离对话语排序的能力,因为恒定偏差或错误的输出尺度会被校准本身精确消除。在这一点与从少量样本估计每个系数的成本之间进行权衡,得出一个标准,该标准解释了哪个映射在哪个语料库和哪个标注预算下胜出,同时提供一种无需硬性决策的收缩变体。我们的发现建议通过线性相关性与真实距离来选择合成检查点,而不是通过绝对误差。代码、数据集和分析可在以下网址获取:https://this https URL。

英文摘要

Speaker distance estimators are trained almost exclusively on simulated room acoustics, because real recordings annotated with the true talker-to-microphone distance are scarce. We show that models trained this way transfer poorly. On three real corpora we evaluate, simply predicting the average distance of the corpus is more accurate than any learned model. Then, we ask how few labelled real utterances are needed to make a frozen, synthetic-trained estimator useful, and study post-hoc calibration maps that rescale its output without gradients or retraining. An analysis of the achievable error shows that what the calibration is not limited by the absolute accuracy of the estimator, but how well it orders utterances by distance, since a constant bias or a wrong output scale is removed exactly by the calibration itself. Balancing this against the cost of estimating each coefficient from few samples yields a criterion that accounts for which map wins on which corpus and at which annotation budget, together with a shrinkage variant that requires no hard decision. Our findings suggest selecting synthetic checkpoints by linear correlation with true distances rather than by absolute error. Code, datasets, and analysis are available at https://github.com/michaelneri/audio-distance-estimation.

CommentsSubmitted to IEEE International Conference of Acoustics, Speech, and Signal Processing (IEEE ICASSP 2027)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑