发表机构
Hospital Beatriz Ângelo, Faculdade de Medicina Universidade Católica Portuguesa(葡萄牙天主教大学医学院贝阿特丽斯·安热洛医院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出端到端模型DiaWhisper及失败挖掘的DPO优化DiaWhisper-DPO,用于临床访谈转录与角色归属,在DAIC-WOZ和PDCH-HAMD上显著提升准确率并降低错误率。
AI 中文摘要
从临床访谈中进行自动化抑郁筛查需要将话语归属于临床医生或患者。我们评估了两个数据集:DAIC-WOZ,其中仅包含参与者的录音,需要重新合成双方以进行受控的双人评估;以及PDCH-HAMD,包含语音转换的真实中文访谈,用于跨语言验证。级联系统将说话人分离与角色分配启发式方法相结合,因此错误可能在各个阶段传播。我们提出了一种端到端模型,命名为DiaWhisper,该模型使用LoRA和辅助帧级角色头对Whisper-large-v3进行微调,用于转录和归属,同时提出DiaWhisper-DPO,一种失败挖掘的优化方法,利用真实的解码失败作为DPO的拒绝完成,无需人工偏好标注。在29个DAIC-WOZ测试会话中,DiaWhisper-DPO实现了0.973的角色准确率和0.119的DER,比最强的级联基线低72%,并将种子变异从σ=.205降低到.002。在PDCH-HAMD上重新训练后,它实现了0.757的角色准确率,并改善了所有78个会话-种子对。
英文摘要
Automated depression screening from clinical interviews requires attribution of utterances to the clinician or patient. We evaluate two datasets: DAIC-WOZ, where participant-only recordings require re-synthesizing both sides for controlled two-party evaluation, and PDCH-HAMD, comprising voice-converted real Chinese interviews for cross-lingual validation. Cascaded systems combine speaker diarization with role-assignment heuristics, so errors can propagate across stages. We propose an end-to-end model, which we named DiaWhisper, that fine-tunes Whisper-large-v3 with LoRA and an auxiliary frame-level role head for transcription and attribution, together with DiaWhisper-DPO, a failure-mined refinement that uses genuine decoding failures as DPO rejected completions without human preference annotation. On 29 DAIC-WOZ test sessions, DiaWhisper-DPO achieves 0.973 role accuracy and 0.119 DER, 72% below the strongest cascaded baseline, and reduces seed variation from σ = .205 to .002. Retrained on PDCH-HAMD, it achieves 0.757 role accuracy and improves all 78 session-seed pairs.
Comments5 pages, 2 figures. Submitted to ICASSP 2027