arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

角色引导的临床精神病学语音录音说话人删除验证与音频语言模型

Role-guided Speaker Deletion Verification in Clinical Psychiatry Speech Recordings with Audio Language Models

Joseph T Colonel, Daniel Katzman, Kelsey Kirker, Adam N Davidson, Shalaila S Haas, Cheryl Corcoran, René S Kahn, Guillermo Checci, Baihan Lin

arXiv 2609.38491首次发表:更新:

发表机构

Icahn School of Medicine at Mount Sinai(西奈山伊坎医学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对精神病学临床对话录音中按角色删除说话人的验证问题,提出双流水线方法,利用音频语言模型和大语言模型扫描残留音频,并在48个二元录音上评估,OR集成F1达0.478,凸显模型互补性。

AI 中文摘要

精神病学临床研究日益依赖于大规模收集口语数据来识别声学和语言生物标志物。然而,不断变化的知情同意和协议要求可能迫使研究者从多说话人录音中删除指定说话人,并在无法人工审查整个语料库的规模上验证该删除操作。我们针对精神病学中角色驱动的二元临床对话研究这一验证问题,并通过两条平行、对称的流水线对其进行调查:确认临床医生语音已从精神病访谈录音中删除,以及确认患者语音已从相同录音中删除。每条流水线对其目标角色进行原始音频的删减,然后使用音频语言模型和大语言模型扫描剩余输出以识别遗漏的删除。我们在从精神病学环境中提取的48个二元录音语料库上评估了该方法,在仅推理设置下测试了四个开放权重模型:Gemma-4-12B、Gemma-4-31B、Nemotron-3-Nano和Nemotron-3-Nano-Omni。一个由十四个模型视图配置组成的析取OR集成实现了0.478的综合F1分数(精确率0.330,召回率0.870),相比单个模型估计有所改进,这得益于召回率的提升,表明模型和上下文视图之间存在显著的互补性。

英文摘要

Clinical research in psychiatry increasingly relies on large scale collection of spoken language data to identify acoustic and linguistic biomarkers. Yet evolving consent and protocol requirements can oblige investigators to remove a designated speaker from multi-speaker recordings and to verify said removal at a scale infeasible for manual review of entire corpora. We study this verification problem for role-driven dyadic clinical dialogue in psychiatry and investigate it with two parallel, symmetric pipelines: confirming that clinician speech has been removed from psychiatric interview recordings, and confirming that patient speech has been removed from the same recordings. Each pipeline redacts the raw audio for its target role and then scans the surviving output with audio-language and large-language models to identify missed deletions. We evaluate this approach on a corpus of 48 dyadic recordings drawn from psychiatry settings, testing four open-weight models in an inference-only setting: Gemma-4-12B, Gemma-4-31B, Nemotron-3-Nano, and Nemotron-3-Nano-Omni. A disjunctive OR ensemble over fourteen model-view configurations had a combined F1 of 0.478 (precision 0.330, recall 0.870), an improvement over individual model estimates driven by recall gains that point to substantial complementarity across models and context views.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑