医患对话的稳健摘要:超越转录挑战的塔尔图大学系统
Robust Summarization of Doctor-Patient Conversations: TalTech Systems for the Beyond Transcription Challenge
浏览论文内容
中文总结 AI 辅助
塔尔图大学针对超越转录挑战,筛选鲁棒语音语言模型,用LoRA微调及DAPO强化学习调整Voxtral模型。其系统在两轨道夺冠,幻觉率最低,还表明微调文本转录本能提升语音输入鲁棒性。
中文摘要 AI 辅助
本文介绍了塔尔图大学提交给超越转录挑战(BeTraC)的内容,该挑战要求直接从医患长对话录音中生成SOAP记录,无需中间转录。在筛选了对长音频具有鲁棒性的开放权重语音语言模型后,我们通过LoRA监督微调,然后使用挑战指标Open Medical Concept F1作为奖励的DAPO强化学习,对Voxtral Mini(轻量级轨道)和Voxtral Small(重量级轨道)进行了调整。我们的系统在两个轨道上均排名第一,独立的基于语言模型的评判评估显示,在所有提交的内容中,我们的系统幻觉率最低,这表明针对概念匹配指标的强化学习不必牺牲事实可靠性。我们还发现,对文本转录本进行微调可以很好地转移到语音输入上,并且似乎可以提高对域外真实录音的鲁棒性。
英文摘要
This paper describes TalTech's submissions to the Beyond Transcription Challenge (BeTraC), which requires generating SOAP notes directly from long doctor-patient conversation recordings, without intermediate transcription. After screening open-weight speech LLMs for long-audio robustness, we adapted Voxtral Mini (lightweight track) and Voxtral Small (heavyweight track) with LoRA supervised fine-tuning followed by DAPO reinforcement learning that uses the challenge metric, Open Medical Concept F1, as its reward. Our systems ranked first in both tracks, and an independent LLM-as-a-judge evaluation showed the lowest hallucination rate among all submissions, indicating that reinforcement learning against a concept-matching metric need not compromise factual reliability. We also find that fine-tuning on text transcripts transfers well to speech input and appears to improve robustness on out-of-domain real recordings.
发表机构
- TalTech(塔林理工大学)
机构由 AI 辅助整理,请以论文原文为准。