arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

文本生成交响曲:临床笔记生成的基准测试

Symphony for Text Generation: Benchmarking Clinical Note Generation

Daniel Varab, Victor Petrén Bach Hansen, Asbjørn W. Helge, Kevin Pelgrims, Mathias Baltzersen, Adrian Young-San Roessler, Vanessa Klungtvedt, Maximilian Brand, Lasse Krogsbøll, Henrik Cullen, Lars Maaløe

arXiv 2610.08161首次发表:更新:

AI 中文总结

本文提出MedConv数据集和受控评估框架,证明Corti的API文本生成在临床笔记质量上不逊于甚至优于商业抄写软件,并支持按需微调。

AI 中文摘要

环境文档系统正迅速获得采用,但其对临床笔记质量的影响仍缺乏充分表征。我们引入了MedConv,一个包含英语、丹麦语和德语共300次临床就诊的多语言数据集,并将其与环境临床智能基准(ACI-BENCH)一同用于比较临床AI平台Corti与两款基于通用AI构建的主流且可访问的环境抄写软件应用。我们提出了一个受控临床评估框架,该框架结合了蕴含指标与LLM评判的成对比较,涵盖从PDSQI-9采用的八个维度。结果表明,Corti基于API的文本生成基础设施与领先的商业抄写软件相当或更优。我们进一步表明,Corti可配置的API提供了必要的灵活性,以针对特定文档使用场景微调质量维度。我们展示了评估方法并发布了一个数据集,以支持未来对环境文档系统的可重复比较。

英文摘要

Ambient documentation systems are rapidly gaining adoption, yet their impact on clinical note quality remains poorly characterized. We introduce MedConv, a multilingual dataset of 300 clinical encounters in English, Danish, and German, and use it alongside the Ambient Clinical Intelligence benchmark (ACI-BENCH) to compare Corti, a clinical AI platform, with two leading, accessible ambient scribe software applications built on general-purpose AI. We present a controlled clinical evaluation framework that combines entailment metrics with LLM-judged pairwise comparisons across eight dimensions adopted from PDSQI-9. Results show that Corti's API-based text-generation infrastructure is on par with or outperforms leading commercial scribes. We further show that Corti's configurable API provides the flexibility necessary to fine-tune quality dimensions for specific documentation use cases. We present the evaluation methodology and release a dataset to support future reproducible comparison of ambient documentation systems.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑