发表机构
Hippocratic AI(希波克拉底人工智能公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出面向对话式医疗智能体的HealthCUES系统,可实时检测分析咳嗽等呼吸信号,经内部与外部数据集评估及医疗人员验证,性能优异且具临床实用性。
AI 中文摘要
实时口语对话中的咳嗽事件蕴含具有临床价值的呼吸信号,但现有对话系统将其视为需丢弃的声学噪声。我们提出HealthCUES(基于实体声音的临床理解),这是一种用于实时对话式智能体副语言呼吸监测的流式管道,据我们所知,所有现有系统均不具备该能力。HealthCUES通过与对话轮次边界对齐的滚动缓冲处理音频,实现亚秒级事件检测且不中断对话流。除二元咳嗽检测外,该系统提供细粒度分析:(i)区分咳嗽与清嗓子动作;(ii)咳嗽亚型分类(干咳、湿咳、犬吠样咳、鸡鸣样咳)并给出置信度分数;(iii)带起止边界的时长估计。为避免警报疲劳,HealthCUES引入对话感知门控机制,基于对话上下文调节触发条件。该系统利用多模态大语言模型(MLLM)Qwen3Omni,通过约束结构化输出将咳嗽分析分解为并行预测任务以实现独立提示优化。对847段内部对话音频片段的评估显示,咳嗽检测F1值为93%,湿咳/干咳亚型分类加权F1值为0.75,平均端到端延迟为340ms;在AMI会议语料库上的外部验证确认,在语音存在情况下可实现稳健的咳嗽、清嗓子与语音分离(宏F1值为0.91)。一项由持牌医疗专业人员参与的用户研究证实,亚型信息具有临床相关性,且该系统在远程医疗工作流程中具有实用性。
英文摘要
Cough events during live spoken conversations carry clinically valuable respiratory signals, yet existing dialogue systems treat them as acoustic noise to be discarded. We present HealthCUES (Clinical Understanding from Embodied Sounds), a streaming pipeline for paralinguistic respiratory monitoring in real-time conversational agents, a capability that, to the best of our knowledge, is absent from all prior systems. HealthCUES processes audio through a rolling buffer aligned with dialogue turn boundaries, enabling sub-second event detection without interrupting conversational flow. Beyond binary cough detection, the system provides fine-grained analytics: (i) differentiation between coughing and throat clearing, (ii) cough subtype classification (dry, wet, barking, whooping) with confidence scores, and (iii) temporal duration estimation with start-end boundaries. To prevent alert fatigue, HealthCUES introduces dialogue-aware gating mechanisms that modulate triggering based on conversational context. The system leverages Qwen3Omni, a multimodal large language model (MLLM), with constrained structured outputs, decomposing cough analysis into parallel prediction tasks for independent prompt optimization. Evaluation on 847 in-house conversational audio segments demonstrates 93\% F1 for cough detection, 0.75 weighted-F1 for wet/dry subtype classification, and average end-to-end latency of 340ms; external validation on the AMI meeting corpus confirms robust cough, throat-clearing, and speech separation in the presence of speech (0.91 macro-F1). A user study with licensed healthcare professionals confirms the clinical relevance of subtype information and the system's utility in telehealth workflows.
CommentsAccepted for publication at SIGDIAL 2026