超越转录的推理:音频语言模型在儿童口吃语音上的应用
Reasoning Beyond Transcription: Audio Language Models on Child Stuttering Speech
浏览论文内容
中文总结 AI 辅助
本研究探索音频语言模型在混合说话者环境中对儿童口吃语音的推理能力,通过语义摘要和语音蕴含任务,发现模型能提取高层次含义但推理随不流畅程度增加而显著下降。
中文摘要 AI 辅助
儿童语音在声学、韵律和语言结构上与成人语音不同。言语不流畅(如重复)进一步挑战了自动理解。虽然音频语言模型(ALMs)在从语音音频中进行强语义推理方面表现出色,但它们在混合说话者环境中对不流畅儿童语音进行推理的能力仍未得到探索。我们通过两个任务对此进行研究:以儿童为中心的语义摘要和语音蕴含。实验使用了在混合说话者访谈中口吃儿童的录音,没有明确的说话者分离。模型被指令引导以关注儿童,保留临床相关的不流畅,并避免成人语音泄漏。评估结合了基于LLM的评判和基于参考的指标,并以转录预言机基线为锚点来隔离错误。结果表明,虽然ALMs能从口吃语音中提取高层次含义,但随着不流畅程度的增加,推理能力显著下降。
英文摘要
Child speech differs from adult speech in acoustics, prosody, and linguistic structures. Speech disfluencies (such as repetitions) further challenge automatic understanding. While Audio Language Models (ALMs) show strong semantic reasoning from speech audio, their ability to reason about disfluent child speech in mixed-speaker settings remains unexplored. We investigate this through two tasks: child-focused semantic summarization and speech entailment. Experiments use recordings of children who stutter in mixed speaker interviews without explicit speaker separation. Models are instruction-guided to focus on the child, preserve clinically relevant disfluencies, and avoid adult-speech leakage. Evaluation combines LLM-based judges and reference-based metrics, anchored by transcript-oracle baselines to isolate errors. Results show that while ALMs extract high-level meaning from stuttered speech, reasoning degrades significantly with increased
发表机构
- University of Florida(佛罗里达大学)
机构由 AI 辅助整理,请以论文原文为准。