发表机构
Southeast University; School of Biological Science & Medical Engineering, Southeast University(东南大学; 东南大学生物科学与医学工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对构音障碍患者在专业演讲场景的沟通问题,提出基于三级级联ASR-LLM-TTS架构的Re-Sonance系统,集成多种技术,经评估其能提高轻度至中度构音障碍患者语音清晰度与自然度,验证了基于LLM方法增强语音驱动AAC系统的潜力。
AI 中文摘要
构音障碍患者在会议等专业演讲场景中面临实时沟通挑战。现有辅助和替代沟通(AAC)系统因高延迟和不自然语音模式无法满足需求。本文提出Re-Sonance,一种用于实时专业演讲场景的新型基于大语言模型(LLM)增强的语音驱动AAC系统。通过集成Whisper ASR、Qwen LLM和CosyVoice TTS,它在保持实时性能的同时提高了语音清晰度和自然度。使用普通话构音障碍语音数据集的主客观评估表明,该语音重建方法显著提高了轻度至中度构音障碍患者的清晰度并保持语义连贯。虽对严重构音障碍效果有限,但验证了基于LLM方法增强语音驱动AAC系统的潜力。
英文摘要
Individuals with dysarthria face significant challenges in professional speaking scenarios such as conferences, presentations, and meetings, where real-time communication is crucial. While existing Augmentative and Alternative Communication (AAC) systems provide basic support, they often fail to meet the demands of professional speaking environments due to high latency and unnatural speech patterns. This paper presents Re-Sonance, a novel LLM-enhanced speech-driven AAC system designed for real-time professional speaking scenarios. By integrating Whisper ASR, Qwen LLM, and CosyVoice TTS, Re-Sonance achieves improved speech intelligibility and naturalness while maintaining real-time performance. Both subjective and objective evaluations using a Mandarin dysarthric speech dataset demonstrate that our speech reconstruction approach significantly improved intelligibility while preserving semantic coherence for speakers with mild to moderate dysarthria. Although performance remains limited for severe dysarthria cases, our findings validate the potential of LLM-based methods for enhancing speech-driven AAC systems, paving the way for more effective and accessible communication technologies.
CommentsAccepted by NCMMSC 2025