arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

针对构音障碍和气管造口说话人的个性化自动语音识别:基于人工对话方法

Personalized Automatic Speech Recognition for a Dysarthric and Tracheostomic Speaker using Artificial Conversations

David Nadrchal, Monorama Swain, Florian Schmid, Gerhard Widmer, Paul Primus

arXiv 2610.03017首次发表:更新:

AI 中文总结

针对一位捷克语气管造口及严重构音障碍说话人,本文提出基于Whisper Base的多阶段微调ASR系统,利用人工对话协议收集的33小时数据,实现字符错误率相对降低50%,证明严重受损语音也可获得有效识别。

AI 中文摘要

本工作提出了一种为一位捷克语说话人定制的自动语音识别(ASR)系统,该说话人因永久性气管造口和严重构音障碍,其言语对未经训练的听者而言难以理解。我们发布了一个公开数据集,包含该说话人33小时的带标注语音,这些数据采用一种新颖的“人工对话”协议收集,该协议旨在实现高参与度和对话真实性。我们提出了一种基于Whisper Base的多阶段训练流程:在标准捷克语语音、声学模拟的气管造口语音以及该说话人的数据上依次进行微调。我们在三种近实时场景下评估该系统:脚本对话、问答和自发对话,与Whisper Base基线相比,字符错误率相对降低了50%,并在孤立话语的声学识别中超过了其助手的平均识别准确率。我们证明,即使对于严重受损的语音,也能实现有用的ASR,这由定量结果和说话人的反馈所证实。

英文摘要

This work presents an automatic speech recognition (ASR) system personalized for a Czech speaker with a permanent tracheal stoma and severe dysarthria rendering their speech unintelligible to untrained listeners. We release a public dataset containing 33 annotated hours of the speaker's speech, collected using a novel "artificial conversation" protocol designed for high engagement and dialogue realism. We propose a multi-stage training pipeline based on Whisper Base: fine-tuning on standard Czech speech, acoustically simulated tracheostomic speech, and the speaker's data. We evaluate the system across three near real-time scenarios: scripted conversations, question answering, and spontaneous dialogue, achieving a 50\% relative reduction in Character Error Rate compared to Whisper Base baseline and surpassing the average recognition accuracy of their assistants in acoustic recognition of isolated utterances. We demonstrate that even for severely impeded speech, a helpful ASR is achievable, as evidenced by the quantitative results and the feedback from the speaker.

Comments8 pages, three figures, to be published in IEEE Speech Language Technology workshop 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑