发表机构
Tianjin University; Shenzhen Institutes of Advanced Technology, Chinese Academy of Sciences; Fuzhou University; Huiyan Technology (Tianjin) Co., Ltd(天津大学; 中国科学院深圳先进技术研究院; 福州大学; 慧眼科技(天津)有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对阿尔茨海默病语音检测,提出LLM锚定的副语言增强方法,通过韵律事件文本化、词汇-韵律单元化和文本锚定融合,在ADReSS和ADReSSo上取得最先进性能。
AI 中文摘要
基于语音的阿尔茨海默病(AD)自动检测为早期认知筛查提供了一种非侵入性且可扩展的方法。AD影响词汇语义组织和语音产生,包括非典型停顿和单词延长。然而,现有方法尚未完全将这些副语言线索与语言内容整合。我们提出了LLM锚定的副语言增强(LAPE),通过三项协同创新,用副语言线索丰富LLM衍生的语言表征。第一项是韵律事件文本化,通过将停顿和延长编码为具有有界持续时间感知重复的显式标记,使LLM能够将停顿和延长与词汇内容联合建模。第二项是词汇-韵律单元化和分块,通过仅池化连续词单元,在两种模态中保留事件身份和幅度。第三项是文本锚定的副语言融合,通过使用NormGate对局部和话语级语音特征进行归一化并相对于文本动态缩放,从而整合这些特征。我们在ADReSS和ADReSSo上使用参与者级交叉验证和留一受试者评估来评估LAPE。LAPE在所有四个主要设置中均达到最先进性能。代码将在录用后发布。
英文摘要
Speech-based automatic detection of Alzheimer's disease (AD) provides a non-invasive and scalable approach to early cognitive screening. AD affects both lexical-semantic organization and speech production, including atypical pauses and word elongations. However, existing methods have yet to fully integrate these paralinguistic cues with linguistic content. We propose LLM-Anchored Paralinguistic Enrichment (LAPE), which enriches LLM-derived linguistic representations with paralinguistic cues through three coordinated innovations. The first is prosodic event textualization, which enables the LLM to model pauses and elongations jointly with lexical content by encoding them as explicit markers with bounded duration-aware repetition. The second is lexico-prosodic unitization and chunking, which preserves event identity and magnitude in both modalities by pooling only consecutive word units. The third is text-anchored paralinguistic fusion, which integrates local and utterance-level speech features by using NormGate to normalize and dynamically scale them relative to text. We evaluate LAPE on ADReSS and ADReSSo using participant-level cross-validation and leave-one-subject-out evaluation. LAPE achieves state-of-the-art performance across all four primary settings. Code will be released upon acceptance.
Commentsv2: 13 pages including references and supplementary material, 3 figures, 5 main tables, 8 supplementary tables. This version adds the supplementary material omitted in v1. (v1: 9 pages including references, 3 figures.)