发表机构
Pattern Recognition Lab, Friedrich-Alexander-Universität Erlangen-Nürnberg; GITA Lab. Facultad de Ingeniería. Universidad de Antioquia UdeA; Chair for AI in Healthcare and Medicine, Technical University of Munich(弗里德里希-亚历山大-埃尔朗根-纽伦堡大学模式识别实验室; 安蒂奥基亚大学工程学院GITA实验室; 慕尼黑工业大学医疗与医学人工智能教席)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将PhonoQ-2.0适配至儿童语音,比较多种对齐监督与初始化策略,在58名儿童辅音上显著提升浊音识别,并验证了结构化表征在儿童语音分析中的可解释性与临床适用性。
AI 中文摘要
结构化音系表征为通用语音嵌入提供了一种可解释的替代方案,但现有模型大多基于成人语音训练。我们使用CHILDES-Aligned数据将PhonoQ-2.0适配到儿童语音,并在两种初始化策略(Adult PhonoQ和scratch)下比较了三种对齐监督条件(Adult、Adult+Child和Child-only)。泛化性能通过人工儿童语音标注进行评估。在来自58名典型发育儿童的1,352个辅音目标上,儿童语音适配在所有监督条件下均提升了浊音识别性能,从Adult PhonoQ的0.922宏F1提升至适配后的0.972--0.987。发音方式对对齐监督更为敏感:Adult+Child MFA达到0.804和0.796,而Adult MFA监督下约为0.70。发音部位在所有系统中保持相对较强的性能(0.871--0.902),但各类别表现差异较大。软腭前移对立在所有七个模型变体中均得以保留。纵向UltraPhonix分析进一步揭示了说话者特定的软腭和后齿槽变化,这些变化在模型中基本保留,并与报告的临床进展大体一致。
英文摘要
Structured phonological representations provide an interpretable alternative to generic speech embeddings, but existing models are largely trained on adult speech. We adapt PhonoQ-2.0 to child speech using CHILDES-Aligned data and compare three alignment-supervision conditions (Adult, Adult+Child, and Child-only) across two initialization strategies (Adult PhonoQ and scratch). Generalization is evaluated against manual child-speech annotations. On 1,352 consonant targets from 58 typically developing children, child-speech adaptation improves voicing recognition across all supervision conditions, from 0.922 macro-F1 for Adult PhonoQ to 0.972--0.987 after adaptation. Manner is more sensitive to alignment supervision: Adult+Child MFA reaches 0.804 and 0.796, compared to approximately 0.70 under Adult MFA supervision. Place remains comparatively strong across systems (0.871--0.902), although per-class performance varies substantially. The velar-fronting contrast is preserved across all seven model variants. Longitudinal UltraPhonix analysis further reveals speaker-specific velar and post-alveolar changes that are largely preserved across models and broadly consistent with reported clinical progress.
CommentsSubmitted for review at ICASSP 2027