arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

临床AI系统、医生与前沿语言模型在初级保健诊断中的表现

Performance of Clinical AI System and Physicians and Frontier Language Models in primary care diagnostics

Andy Nkansah, Hanna Plotnitskaya, Stanislau Salavei, Anna Kozlova, Piotr Gibas, Julian Milek, Viktar Harbachou, Aleksey Ropan, Pavel Satalkin

arXiv 2609.09070首次发表:更新:

发表机构

A.I. Doctor Medical Assist LTD(A.I. Doctor Medical Assist有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究在150个合成初级保健咨询中比较了临床AI系统Doctorina、医生和前沿语言模型,发现Doctorina在诊断、检查及治疗评分上均优于医生,且结果可重复。

AI 中文摘要

临床AI评估应涵盖在适应性信息收集后的诊断与管理。我们在150次合成的波兰语初级保健咨询中,比较了Doctorina、八名医生以及四个独立的尖端语言模型。Doctorina的Top-1一致性达到82.0%,而医生为57.0%(差异为25.0个百分点;95%置信区间为17.7-32.7);在主要或参考鉴别诊断一致性上,Doctorina为97.3%,医生为85.0%。在149对病例中,规范化检查与治疗评分分别为89.4对66.9和83.7对61.2。Doctorina在所有六个组别中拥有最高的诊断点估计值;Kimi K3紧随其后,而Claude Opus 5在Opus、Doctorina和Kimi之间紧密排列的管理估计中领先。第二次执行Doctorina在所有结果上重现了对医生的优势。因此,Doctorina相对于医生的优势从初步诊断选择延伸至适应性咨询后更高质量的诊断检查与初始治疗。

英文摘要

Clinical AI evaluation should encompass diagnosis and management after adaptive information gathering. We compared Doctorina, eight physicians and four standalone frontier language models in 150 synthetic Polish-language primary-care consultations. Doctorina achieved 82.0% Top-1 concordance versus 57.0% for physicians (difference, 25.0 percentage points; 95% confidence interval, 17.7-32.7) and 97.3% versus 85.0% primary-or-reference-differential concordance. Across 149 case pairs, normalized workup and treatment scores were 89.4 versus 66.9 and 83.7 versus 61.2. Doctorina had the highest diagnostic point estimates among all six groups; Kimi K3 ranked next, while Claude Opus 5 led the closely spaced management estimates of Opus, Doctorina and Kimi. A second Doctorina execution reproduced the advantages over physicians across all outcomes. Doctorina's advantage over physicians therefore extended from primary-diagnosis selection to higher-rated diagnostic workup and initial treatment after adaptive consultation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑