发表机构
Ashoka University; Microsoft Research India; IIT Madras(阿肖卡大学; 微软印度研究院; 马德拉斯理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对印度多语言、嘈杂的真实临床环境,研究指出当前ACS缺乏公开真实基准,呼吁构建多语言真实世界评估基准以保障安全可靠。
AI 中文摘要
环境临床记录员(ACS)正迅速在全球南方医疗环境中大规模部署,旨在减少临床医生的文档记录时间,尤其是在印度等负担过重的环境中。这些ACS主要基于在全球北方语音、语言和咨询风格上构建和验证的模型开发或提炼而来。印度的临床接触具有简短、三方、多语言、与低资源语言混杂以及在高资源受限、嘈杂环境中进行的特点——这多倍增加了自动语音识别(ASR)和笔记生成错误的发生概率。我们主张迫切需要开发一个标准化的评估基础设施,以评估这些系统在印度医疗环境中是否安全、可靠且适用。我们通过一项混合方法研究来证实我们的主张——包括对公开可用的患者-临床医生对话数据集进行系统性调查,将这些数据集与源自印度临床沟通文献中的对话和文化标记进行定量比较,以及对在印度和非洲构建和部署ACS的五家组织进行半结构化访谈。我们的调查显示,在印度没有公开可用的大规模真实世界ACS基准,现有数据集绝大多数是合成的。我们注意到,可用的全球北方数据集与印度临床接触的预期对话和文化结构存在显著差异。最后,我们的访谈揭示,部署组织各自构建了专有的、不可比较的评估流程,形成了一个碎片化的生态系统,缺乏独立可靠的采购依据。我们呼吁开发一个公开共享的、真实世界的、多语言的ACS评估基准,并概述了这样一个基准所需具备的特性和政策。
英文摘要
Ambient clinical scribes (ACS) are being rapidly deployed at scale across Global South healthcare settings, aiming to reduce clinician documentation time, especially in overburdened environments like India. These ACS are primarily developed or distilled from models built and validated on Global North speech, languages and consultation styles. Indian clinical encounters are brief, triadic, multilingual, code-mixed with low-resource languages, and conducted in highly resource-constrained, noisy settings -- increasing the likelihood of ASR and note-generation errors manyfold. We posit an urgent need to develop a standardized evaluation infrastructure to assess whether these systems are safe, reliable, and well-suited to the Indian healthcare setting. We substantiate our claims through a mixed-methods study -- a systematic survey of publicly available patient-clinician conversational datasets, a quantitative comparison of these datasets against conversational and cultural markers drawn from the Indian clinical-communication literature, and semi-structured interviews with five organizations building and deploying ACS in India and Africa. Our survey shows that there are no publicly available, large-scale, real-world benchmarks for ACS in India, with existing datasets being overwhelmingly synthetic. We note that the available Global North datasets diverge significantly from the expected conversational and cultural structures of Indian encounters. Finally, our interviews reveal that deploying organizations have each built proprietary, incomparable evaluation pipelines, creating a fragmented ecosystem with no independent and reliable basis for procurement. We call for the development of a publicly shared, real-world, multilingual benchmark for ACS evaluation and outline the properties and policies such a benchmark would require.
CommentsUnder Submission