arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当患者打断时:将临床对话人工智能的安全性扩展至打断场景

When Patients Cut In: Extending Clinical Conversational AI Safety to Interruptions

Zachary Ellis, Spencer Hazel, Adam Brandt, Yajie Vera He, Ernest Lim, Jared Joselowitz

arXiv 2608.29241首次发表:更新:

发表机构

Ufonia Limited; Newcastle University(乌福尼亚有限公司; 纽卡斯尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对临床对话AI未考虑患者打断的问题,调整重叠类别为三种操作类型,测试四种LLM配置,发现竞争型FAQ打断下所有模型均出现信息覆盖失败,提出需按单元评估打断鲁棒性。

AI 中文摘要

临床语音智能体现已部署于日常医疗护理中,而真实患者不会按顺序等待,他们会进行打断。这类系统通常采用级联架构(语音转文字 -> 大语言模型(LLM) -> 文字转语音),因此当患者在智能体发声中途打断时,即便模型能处理协作式转录内容,临床所需信息仍可能丢失。然而,临床对话人工智能基准几乎普遍假设患者会等待智能体完成发言,忽略了打断导致的所需信息丢失问题。我们针对打断恢复开展基于转录文本的评估,将会话分析中的重叠类别调整为三种操作类型(识别型、竞争型、过渡子单元型),并在涵盖问诊(信息收集)和常见问题解答(信息提供)的四个单元中测试四种面向部署的非推理大语言模型配置,评估指标为智能体是否保留临床所需内容。在信息收集单元中,不同模型的目标问题失败情况存在差异;在可直接比较的信息提供单元中,所有模型的失败率均上升。不同单元的排名存在差异,竞争型常见问题解答打断场景下,所有四个模型的30个信息提供覆盖案例均失败(威尔逊95%置信区间:88.6%-100.0%;三个模型的基准为0/30,Llama模型的基准为4/30)。简短的道歉标记(“抱歉打断一下”)会使恢复率产生数十个百分点的变化,且不同模型的表现不一致,其中一个模型的恢复率反而降低。因此,打断鲁棒性不能用单一分数衡量:评估必须以内容为基础,按单元报告,并与部署场景的打断特征相匹配。

英文摘要

Clinical voice agents are now deployed in routine care, where real patients do not wait their turn: they interrupt. These systems typically use a cascaded architecture (speech-to-text -> LLM -> text-to-speech), so when a patient cuts the agent off mid-utterance, clinically required content can be lost even when the model handles cooperative transcripts well. Yet clinical conversational-AI benchmarks almost universally assume patients wait for the agent to finish, missing interruption-induced loss of required content. We present a transcript-based evaluation of interruption recovery, adapting conversation-analytic overlap categories into three operational types (recognitional, competitive, transitional sub-unit) and testing four deployment-oriented, non-reasoning LLM configurations across four cells spanning history-taking (information gathering) and FAQ (information provision), scored on whether the agent preserves the clinically required content. In the gathering cells, target-question failure varied across models; in the provision cells, where arms are directly comparable, failure rose for every model. Rankings differ across cells, and competitive FAQ interruption produced 30/30 provision-coverage failures for all four models (Wilson 95% CI: 88.6-100.0%; baseline 0/30 for three, 4/30 for Llama). A brief apology marker ("sorry to interrupt") shifts recovery by tens of percentage points, inconsistently across models, and for one it reduces recovery. Interruption robustness therefore cannot be a single score: evaluation must be content-grounded, reported per cell, and matched to the deployment's interruption profile.

CommentsAccepted to Findings of EMNLP 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑