arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.31165cs.SDcs.LGeess.AS

BreathGRU:一种用于呼吸音频中语音与呼吸分割的新型半监督双向门控循环单元框架

BreathGRU: A Novel Semi-Supervised Bidirectional Gated Recurrent Unit Framework for Speech and Breath Segmentation for Respiratory Audio

Sania Fatima Sayed, John W. Holloway, Reyer Zwiggelaar, Faisal I. Rezwan

首次发表
浏览论文内容

中文总结 AI 辅助

针对呼吸音频分析中语音-呼吸分割的不足,提出半监督双向门控循环单元框架BreathGRU,结合特征提取、双向建模、伪标签精炼和维特比解码,在多项指标上优于现有方法,实现高效分割。

中文摘要 AI 辅助

语音-呼吸分割是呼吸音频分析中的一项基础预处理步骤,可支持呼吸声学生物标志物提取、肺功能预测和疾病监测等应用。现有方法,包括阈值方法、基于傅里叶变换的技术以及无监督和预训练的语音活动检测(VAD)模型,主要侧重于语音检测,通常将呼吸事件分类为非语音或静音,限制了其在精确呼吸检测中的适用性。为解决这一局限,我们提出了BreathGRU,一种专门为语音-呼吸分割设计的半监督双向门控循环单元(BiGRU)框架。该框架结合了帧级声学特征提取、双向循环建模、伪标签精炼和持续时间约束的分段维特比解码,以生成语音和呼吸分割结果。我们使用手动标注的录音,将BreathGRU与现有方法进行了评估。性能评估采用了基于事件、基于时间、基于重叠、基于持续时间和基于边界的分割指标。实验结果表明,BreathGRU实现了最高的呼吸事件召回率(0.83)、最低的起始点定位误差(0.14秒)和最高的平均匹配交并比(0.81),并且与大型预训练VAD模型(如Silero)相比,整体分割性能具有竞争力。对手动标注录音的定性评估进一步显示,BreathGRU与手动标注之间具有高度一致性,且呼吸检测优于Silero。这些发现表明,显式的呼吸事件建模相比通用VAD模型具有优势,并确立了BreathGRU作为一种有效的语音-呼吸分割框架,可应用于呼吸音频分析和肺部健康应用。

英文摘要

Speech-breath segmentation is a fundamental preprocessing step in respiratory audio analysis, enabling applications such as respiratory acoustic biomarker extraction, lung function prediction and disease monitoring. Existing approaches, including threshold methods, Fourier Transform-based techniques, and unsupervised and pretrained voice activity detection (VAD) models, primarily focus on speech detection and often classify breathing events as non-speech or silence, limiting their applicability for precise breath detection. To address this limitation, we propose BreathGRU, a semi-supervised Bidirectional Gated Recurrent Unit (BiGRU) framework specifically designed for speech-breath segmentation. The proposed framework combines frame-level acoustic feature extraction with bidirectional recurrent modelling, pseudo-label refinement and duration-constrained Segmental Viterbi decoding to produce speech and breath segmentation. BreathGRU was evaluated against the existing approaches, using manually annotated recordings. Performance was assessed using event-based, time-based, overlap-based, duration-based and boundary-based segmentation metrics. Experiment results demonstrated that BreathGRU achieved the highest breath event recall (0.83), the lowest onset-localisation error (0.14s) and the highest Mean Match Intersection over Union (0.81), with competitive overall segmentation performance compared to large pretrained VAD models like Silero. Qualitative evaluation on manually annotated recordings further showed close agreement between BreathGRU and manual annotation, with better breath detection compared to Silero. These findings demonstrate that explicit breath event modelling provides advantages over general-purpose VAD models and establish BreathGRU as an effective speech-breath segmentation framework which can be applied for respiratory audio analysis and pulmonary healthcare applications.

发表机构

  • Aberystwyth University(阿伯里斯特威斯大学)
  • University of Southampton(南安普顿大学)

机构由 AI 辅助整理,请以论文原文为准。

↑