AI 中文总结
本文提出MedConv数据集和受控评估框架,证明Corti的API文本生成在临床笔记质量上不逊于甚至优于商业抄写软件,并支持按需微调。
AI 中文摘要
环境文档系统正迅速获得采用,但其对临床笔记质量的影响仍缺乏充分表征。我们引入了MedConv,一个包含英语、丹麦语和德语共300次临床就诊的多语言数据集,并将其与环境临床智能基准(ACI-BENCH)一同用于比较临床AI平台Corti与两款基于通用AI构建的主流且可访问的环境抄写软件应用。我们提出了一个受控临床评估框架,该框架结合了蕴含指标与LLM评判的成对比较,涵盖从PDSQI-9采用的八个维度。结果表明,Corti基于API的文本生成基础设施与领先的商业抄写软件相当或更优。我们进一步表明,Corti可配置的API提供了必要的灵活性,以针对特定文档使用场景微调质量维度。我们展示了评估方法并发布了一个数据集,以支持未来对环境文档系统的可重复比较。
英文摘要
Ambient documentation systems are rapidly gaining adoption, yet their impact on clinical note quality remains poorly characterized. We introduce MedConv, a multilingual dataset of 300 clinical encounters in English, Danish, and German, and use it alongside the Ambient Clinical Intelligence benchmark (ACI-BENCH) to compare Corti, a clinical AI platform, with two leading, accessible ambient scribe software applications built on general-purpose AI. We present a controlled clinical evaluation framework that combines entailment metrics with LLM-judged pairwise comparisons across eight dimensions adopted from PDSQI-9. Results show that Corti's API-based text-generation infrastructure is on par with or outperforms leading commercial scribes. We further show that Corti's configurable API provides the flexibility necessary to fine-tune quality dimensions for specific documentation use cases. We present the evaluation methodology and release a dataset to support future reproducible comparison of ambient documentation systems.