arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

利用语音大语言模型表示进行多语言和跨语言帕金森病检测

Exploiting Speech LLM Representations for Multilingual and Cross-Lingual Parkinson's Disease Detection

Sarthak Giri, Zi Haur Pang, Tatsuya Kawahara

arXiv 2609.14431首次发表:更新:

发表机构

Kyoto University(京都大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究语音大语言模型内部表示用于多语言和跨语言帕金森病检测,发现编码器表示优于解码器,并提出基于挤压-激励的动态层聚合框架以提升检测性能。

AI 中文摘要

语音大语言模型(Speech LLMs)已在多种任务中展现出强大性能,但其在病理语音分析中的效用仍未得到充分探索。在本工作中,我们研究了语音大语言模型编码器和解码器组件的内部表示在跨多语言和跨语言环境下用于帕金森病(PD)检测的有效性。我们的研究结果表明,在大多数模型和设置中,编码器表示始终优于解码器对应部分,并且随着音频表示被投影到语言模型空间,病理线索可能逐渐减弱。我们进一步表明,与内部表示相比,生成输出在临床任务中可靠性较低。为了利用分布在多个层中的信息,我们提出了一种基于挤压-激励(SE)的动态层聚合框架,该框架在多个实验中优于最佳层选择,这表明与PD相关的声学线索分布在Transformer层中,而非集中于某一层。

英文摘要

Speech Large Language Models (Speech LLMs) have shown strong performance across diverse tasks, yet their utility for pathological speech analysis remains underexplored. In this work, we investigate the effectiveness of internal representations from encoder and decoder components of Speech LLMs for Parkinson's Disease (PD) detection across multilingual and cross-lingual settings. Our findings reveal that encoder representations consistently outperform their decoder counterparts in most models and settings and that pathological cues may be progressively attenuated as audio representations are projected into the language model space. We further show that generative outputs are less reliable for clinical tasks compared to internal representations. To leverage information spread across multiple layers, we propose a Squeeze-and-Excitation (SE)-based dynamic layer aggregation framework, which surpasses best-layer selection in multiple experiments, suggesting that PD-relevant acoustic cues are distributed across transformer layers rather than concentrated in one.

CommentsAccepted at IEEE Spoken Language Technology Workshop (SLT) 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑