弥合长时临床音频中的模态鸿沟:轻量级与重量级端到端SOAP生成对比研究
Bridging the Modality Gap in Long-Form Clinical Audio: A Comparative Study of Lightweight and Heavyweight End-to-End SOAP Generation
浏览论文内容
中文总结 AI 辅助
本研究提出一种完全端到端的多模态系统,直接从长时临床音频生成SOAP笔记,通过多阶段训练和规模扩展,在BeTraC 2026挑战中超越级联ASR+LLM基线,验证了直接多模态优化的有效性。
中文摘要 AI 辅助
对于现代音频-语言模型而言,从长时医患对话中自动生成临床文档仍具挑战性。尽管级联式ASR系统表现良好,但端到端(E2E)模型在长音频上常面临信息丢失和幻觉问题。针对BeTraC 2026挑战赛,ASLP团队提出了一种完全端到端的多模态系统,可直接从音频生成结构化SOAP笔记,绕过中间转录文本。我们构建了包含141万样本的多任务语料库,并应用了多阶段流水线:领域预训练、监督微调和奖励优化。在轻量级(3B)和重量级(30B)约束下评估架构,结果显示每个训练阶段均逐步提升性能。此外,将规模扩展至30B参数显著提升了概念提取和摘要质量。最终,我们的端到端系统持续优于具有代表性的级联ASR+LLM基线,证明了直接多模态优化在临床文档生成中的有效性。
英文摘要
Automating clinical documentation from long-form doctor-patient conversations remains challenging for modern audio-language models. While cascaded ASR systems perform well, end-to-end (E2E) models often struggle with information loss and hallucinations on extended audio. For the BeTraC 2026 challenge, the ASLP team presents a fully E2E multimodal system that generates structured SOAP notes directly from audio, bypassing intermediate transcripts. We constructed a 1.41-million-sample multi-task corpus and applied a multi-stage pipeline: domain pre-training, supervised fine-tuning, and reward optimization. Evaluating the architecture under both Lightweight (3B) and Heavyweight (30B) constraints reveals that each training stage progressively enhances performance. Furthermore, scaling to 30B parameters substantially boosts concept extraction and summarization quality. Ultimately, our E2E systems consistently outperform representative cascaded ASR+LLM baselines, proving the efficacy of direct multimodal optimization for clinical documentation.
发表机构
- Northwestern Polytechnical University(西北工业大学)
机构由 AI 辅助整理,请以论文原文为准。