增强大型音频语言模型与低级声学特征用于构音障碍语音检测
Augmenting Large Audio Language Models with Low-Level Acoustic Features for Dysarthric Speech Detection
- Idiap Research Institute(伊迪亚普研究所)
- École Polytechnique Fédérale de Lausanne (EPFL)(洛桑联邦理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出微调大型音频语言模型(LALMs)结合低级声学特征与说话者人口统计信息进行构音障碍语音检测,优于深度学习基线,Qwen2-Audio-Instruct达最先进性能。
AI中文摘要:
自动构音障碍语音检测方法可以支持传统的临床诊断,后者依赖于言语和语言病理学家进行的高成本且耗时的评估。现有的自动方法主要依赖深度学习(DL)。最近,大型音频语言模型(LALMs)因其在多种任务中的强大表现而成为有前景的替代方案,但其在构音障碍语音检测中的应用尚未确立。我们提出了一个框架,该框架在语音录音上微调LALMs用于构音障碍语音检测,并结合包含低级声学特征和说话者人口统计信息的文本信息。在两个LALMs上,我们的框架优于基于DL的基线,其中Qwen2-Audio-Instruct达到了最先进的性能。一项消融研究表明,在微调期间纳入声学特征和说话者人口统计信息可提高LALM性能,而仅使用LALMs则表现出仅偶然水平的零样本性能。这些发现为将LALMs适应于构音障碍语音检测建立了一种有效的方法。
英文摘要:
Automatic dysarthric speech detection approaches can support traditional clinical diagnosis, which relies on costly and time-consuming evaluation by a speech and language pathologist. Existing automatic approaches predominantly rely on deep learning (DL). More recently, Large Audio Language Models (LALMs) have emerged as a promising alternative given their strong performance across various tasks, but their application to dysarthric speech detection has not yet been established. We propose a framework that fine-tunes LALMs for dysarthric speech detection on speech recordings combined with textual information comprising low-level acoustic features and speaker demographics. Across two LALMs, our framework outperforms DL-based baselines, with Qwen2-Audio-Instruct achieving state-of-the-art performance. An ablation study shows that incorporating acoustic features and speaker demographics during fine-tuning improves LALM performance, while LALMs alone exhibit only chance-level zero-shot performance. These findings establish an effective approach for adapting LALMs to dysarthric speech detection.