arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.32285eess.AScs.SD

音频预处理对口吃检测的影响:类别特定分析

Audio Preprocessing Effects on Stuttering Detection: A Class-Specific Analysis

Anisha Pattanayak, Hanie Kang, Huang-Cheng Chou, Sudarsana Reddy Kadiri

首次发表
浏览论文内容

中文总结 AI 辅助

本研究分析音频预处理链(去噪、响度归一化、Opus编码、语音活动检测)对口吃检测的影响,发现阻塞类别性能下降最大,阈值调整可恢复大部分F1,但ROC-AUC提升有限。

中文摘要 AI 辅助

音频预处理会影响系统检测口吃的效果。我们研究了在SEP-28k上模拟的去噪、响度归一化、Opus编码和语音活动检测链。我们使用冻结的WavLM Base+特征,并报告基于片段级自助重采样的逐点置信区间。在固定阈值0.5下,该链将阻塞F1从0.638降至0.465,其他四个类别的下降幅度较小。所有五个类别的ROC-AUC均下降。阻塞显示出最大的F1和ROC-AUC损失,而声音重复显示出最大的平均精度损失。在处理后的验证音频上调整阈值可将阻塞F1提升至0.630。在处理后的音频上重新训练并调整阈值,得到0.628。因此,阈值调整解释了观察到的阻塞F1恢复的大部分。它不改变ROC-AUC,重新训练仅将其从0.620提升至0.633,而干净音频上的ROC-AUC为0.724。

英文摘要

Audio preprocessing can affect how well a system detects stuttering. We study a simulated chain of denoising,loudness normalisation, Opus coding, and voice activity detection on SEP-28k. We use frozen WavLM Base+ features and report pointwise confidence intervals from episode-level bootstrap resampling. At a fixed threshold of 0.5, the chain reduces block F1 from 0.638 to 0.465, with smaller decreases for the other four classes. ROC-AUC decreases for all five classes. Blocks show the largest F1 and ROC-AUC losses, while sound repetitions show the largest average precision loss. Tuning the threshold on processed validation audio raises block F1 to 0.630. Retraining on processed audio with threshold tuning gives 0.628. Threshold adjustment therefore accounts for most of the observed block F1 recovery. It does not change ROC-AUC, which retraining raises only from 0.620 to 0.633, compared with 0.724 on clean audio.

发表机构

  • University of Southern California(南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

↑