发表机构
NIHR Biomedical Research Centre: Maudsley; Modality.AI Inc; University of California San Francisco; Child Mind Institute; Harvard University; Massachusetts Institute of Technology (MIT); The Hong Kong Polytechnic University; Arizona State University; INESC-ID; Sword Health; Luxembourg Institute of Health; Friedrich-Alexander-Universitat Erlangen-Nurnberg (FAU); Technical University of Munich (TUM); Lincoln Laboratory, Massachusetts Institute of Technology(NIHR莫兹利生物医学研究中心; Modality.AI公司; 加州大学旧金山分校; 儿童心理研究所; 哈佛大学; 麻省理工学院(MIT); 香港理工大学; 亚利桑那州立大学; 葡萄牙系统与计算机工程研究所; Sword Health公司; 卢森堡卫生研究所; 埃尔朗根-纽伦堡弗里德里希-亚历山大大学(FAU); 慕尼黑工业大学(TUM); 麻省理工学院林肯实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对语音生物标志物因测量定义异质导致的可重复性问题,提出协调统一方案,定义核心测量集,以推动临床采用。
AI 中文摘要
语音和嗓音是多维信号,既能捕捉交际意图,又能反映潜在的生理过程,为健康提供了一种独特的、非侵入性的窗口。分析这些信号有望产生数字生物标志物,这些标志物(i)为研究和临床护理提供可扩展、客观的测量工具,(ii)反映多种疾病(包括神经、精神、呼吸和心血管疾病)的存在或进展。然而,要实现这一前景,该领域必须克服由于数据收集、处理和分析实践异质性而导致的普遍可重复性和可泛化性问题。这种异质性的一个主要来源是基础声学测量本身如何定义和计算。在本文中,我们概述了语音生物标志物发现生命周期中的关键考量,从数据收集到机器学习建模再到临床解释,这些考量对于实现可靠、可重复和临床可转化的结果至关重要。其中最重要的是,协调工作需从共同、精确定义的测量定义开始。因此,作为第一步,我们为跨越呼吸、发声、发音和流畅性的最小临床可解释核心语音测量集提供了定义、生理关联和计算实现。最后,我们讨论了正在进行的标准化工作以及在推进基于语音和嗓音的数字生物标志物采用方面仍然存在的开放挑战。
英文摘要
Speech and voice are multidimensional signals that capture both communicative intent and underlying physiological processes, providing a unique, non-invasive window into health. Analyzing these signals has the potential to yield digital biomarkers that (i) provide scalable, objective measurement tools for research and clinical care and (ii) reflect the presence or progression of diverse conditions, including neurological, psychiatric, respiratory, and cardiovascular disorders. Realizing this promise, however, requires the field to overcome pervasive reproducibility and generalizability issues due to heterogeneous data collection, processing, and analysis practices. A major source of this heterogeneity is how underlying acoustic measures themselves are defined and computed. In this paper, we outline key considerations across the speech biomarker discovery lifecycle, from data collection through machine learning modeling to clinical interpretation, needed to achieve reliable, reproducible, and clinically translatable results. Chief among these is the need for harmonization efforts to start from common, precisely specified measure definitions. As a first step, we therefore provide definitions, physiological correlates, and computational implementations for a minimal, clinically interpretable set of core speech measures spanning respiration, phonation, articulation, and fluency. We close by discussing ongoing standardization efforts and the open challenges that remain in advancing the adoption of speech- and voice-based digital biomarkers.