开源语音情感识别在自然场景中文脊柱门诊咨询中的基准测试:一项初步验证研究
Benchmarking Open-Source Speech Emotion Recognition in Naturalistic Mandarin Spine Clinic Consultations: A Pilot Validation Study
浏览论文内容
中文总结 AI 辅助
本研究在自然场景中文脊柱门诊语音中基准测试三个开源SER模型,发现多数类准确率误导,模型对少数情感检测效果差,表明临床部署需领域适应和更强验证。
中文摘要 AI 辅助
语音情感识别(SER)可能实现临床接触中的被动情感监测,但大多数系统是在表演性实验室语音而非自然场景的中文门诊咨询中验证的。我们针对自然场景脊柱门诊语音,将三个开源SER模型(emotion2vec+、SenseVoice、FunASR)与研究者共识参考进行了基准测试,评估了类别不平衡下的少数状态检测。在一项对前瞻性收集的单中心录音的回顾性分析中,音频经过响度归一化,仅保留患者、家属和临床医生之间的对话。六十五段话语(5-50秒;每位参与者一段;31名患者,34名家属)由六名校准标注者标注为六个类别(快乐、悲伤、恐惧、愤怒、中性、惊讶)。共识采用多数投票,结合Fleiss kappa过滤,并对低一致性片段进行临床医生裁决。指标包括未加权准确率(UA)、宏平均每类准确率、类别和样本级加权准确率(WA)以及带自助法95%置信区间的F1。标签不平衡(中性占58.5%);中位Fleiss kappa为0.230(四分位距0.134-0.519)。SenseVoice和FunASR的UA为61.5%(95%置信区间49.2-73.8%),宏平均每类准确率为87.2%,类别级WA为92.5%,样本级WA为24.6%。emotion2vec+的UA为55.4%(95%置信区间42.5-67.7%),宏平均每类准确率为85.1%,类别级WA为90.6%,样本级WA为24.2%。尽管模型间一致性高(90.8%),所有模型对悲伤、恐惧、愤怒和惊讶的召回率接近零。在这项初步研究中,多数类准确率具有误导性:SER在嘈杂的自然场景参考下对少数情感的检测效果不佳。未经领域适应、多模态建模、更强的参考标准和结果验证,不能从表演语料库基准推断临床部署的成熟度。
英文摘要
Speech emotion recognition (SER) may enable passive affect monitoring in clinical encounters, but most systems are validated on acted laboratory speech rather than naturalistic Mandarin outpatient consultations. We benchmarked three open-source SER models (emotion2vec+, SenseVoice, FunASR) against a researcher-consensus reference in naturalistic spine-clinic speech, assessing minority-state detection under class imbalance. In a retrospective analysis of prospectively collected single-center recordings, audio was loudness-normalized and only conversations among patients, family members, and clinicians were retained. Sixty-five utterances (5-50 s; one per participant; 31 patients, 34 family members) were labeled by six calibrated annotators into six categories (Happy, Sad, Fear, Anger, Neutral, Surprised). Consensus used majority vote with Fleiss kappa filtering and clinician adjudication for low-agreement segments. Metrics included unweighted accuracy (UA), macro-average per-class accuracy, class- and sample-level weighted accuracy (WA), and F1 with bootstrap 95% CIs. Labels were imbalanced (Neutral 58.5%); median Fleiss kappa was 0.230 (IQR 0.134-0.519). SenseVoice and FunASR achieved UA 61.5% (95% CI 49.2-73.8%), macro-average per-class accuracy 87.2%, class-level WA 92.5%, and sample-level WA 24.6%. emotion2vec+ yielded UA 55.4% (95% CI 42.5-67.7%) and macro-average per-class accuracy 85.1%, with class-level WA 90.6% and sample-level WA 24.2%. Despite high inter-model agreement (90.8%), all models had near-zero recall for Sad, Fear, Anger, and Surprised. In this pilot, majority-class accuracy was misleading: SER poorly detected minority emotions against a noisy naturalistic reference. Clinical deployment readiness cannot be inferred from acted-corpus benchmarks without domain adaptation, multimodal modeling, stronger reference standards, and outcome validation.
发表机构
- The LKS Faculty of Medicine, The University of Hong Kong(香港大学李嘉诚医学院)
- The University of Hong Kong-Shenzhen Hospital(香港大学深圳医院)
机构由 AI 辅助整理,请以论文原文为准。