AI 中文总结
研究针对语音健康评估中概念瓶颈框架应用有限的问题,提出基于音频语言模型的语音概念瓶颈框架,经微调ALM提取离散可解释分数用于轻量级分类器,在抑郁症和构音障碍评估任务中表现出色,优于相关基线。
AI 中文摘要
可解释性在临床决策支持中至关重要。概念瓶颈框架通过将输入表示为人类可理解的概念并仅基于这些概念进行预测来提高可解释性。然而,其用于基于语音的健康评估的研究仍然有限。在本研究中,我们提出了一种使用音频语言模型(ALM)的可解释健康评估语音概念瓶颈框架。ALM在语音质量评估数据集上进行微调,以增强其对语音概念的理解,并作为独立的概念提取器,为轻量级下游分类器生成离散、可解释的分数。离散概念分数提供直观解释,轻量级分类器便于事后可解释性分析。抑郁症和构音障碍评估任务的结果表明,所提出的框架可以灵活地使语音概念适应不同的健康状况,并始终优于基于openSMILE和基于自监督语音模型的基线。
英文摘要
Interpretability is critical in clinical decision support. Concept bottleneck frameworks improve it by representing inputs as human-understandable concepts and restricting predictions solely on them. However, research on their use for voice-based health assessment remains limited. In this study, we propose a voice concept bottleneck framework for interpretable health assessment using an audio language model (ALM). The ALM is fine-tuned on a voice quality assessment dataset to enhance its understanding of voice concepts and serves as an independent concept extractor, producing discrete, interpretable scores for a lightweight downstream classifier. The discrete concept scores provide intuitive interpretation, while the lightweight classifier facilitates post-hoc interpretability analyses. Results on depression and dysarthria assessment tasks demonstrate that the proposed framework can flexibly adapt voice concepts to different health conditions and consistently outperforms openSMILE-based and self-supervised speech model-based baselines.