发表机构
Columbia University(哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出Co-Annotator,将专家注视与口述提炼为AOI指导的ViT和生物标志物预填充的VLM,在AMD诊断中联合应用可显著提升效率且不降低准确率。
AI 中文摘要
临床AI常仅优化预测性能,未考虑临床医生的决策过程(包括关注区域和记录内容)。本文提出Co-Annotator,将专家的注视和口述内容提炼为两个指导组件:一是注视对齐的视觉Transformer(Vision Transformer),生成与注视点对齐的感兴趣区域(AOIs);二是受本体约束的视觉语言模型(VLM),可预填充视网膜光学相干断层扫描(OCT)的可编辑生物标志物摘要。我们首先收集专家的注视数据和口述内容(US1)以训练模型,这显著提升了诊断准确率和生物标志物生成质量。随后我们将该系统用于眼科住院医师开展对照住院医师研究(US2),证实两种指导模式均安全且各自具有益处:AOI指导通过指导后延续效应产生持久的感知效率提升,VLM指导使生物标志物文档记录的广度翻倍以上。在两家学术机构的联合部署(US3)中,同时提供两种指导模式产生的效率提升远超单一模式:正确诊断的数量每分钟提升40%,评论编辑时间减少67%,且未损害诊断准确率。值得注意的是,US2中两种模式在指导过程中均未提升效率,因此US3中联合指导模式在指导过程中的效率提升更具显著性。专家蒸馏的多模态指导可同时消除临床工作流的两个不同瓶颈(视觉搜索开销和文档记录负担),且不会损害临床医生已达到的诊断准确率。
英文摘要
Clinical AI often optimizes predictive performance without engaging how clinicians decide where to look and what to write. We present Co-Annotator, which distills expert gaze and dictation into two guidance components: a gaze-aligned Vision Transformer producing fixation-aligned areas of interest (AOIs), and an ontology-bounded vision-language model (VLM) that pre-fills editable biomarker summaries for retinal optical coherence tomography (OCT). We first collect expert gaze and dictations (US1) to train the models, significantly improving diagnostic accuracy and biomarker generation. We then deploy the system with ophthalmology residents: a controlled resident study (US2) confirmed each modality is safe and independently beneficial, with AOI guidance producing lasting perceptual efficiency gains through post-guidance carryover and VLM guidance more than doubling biomarker documentation breadth. In a combined deployment across two academic institutions (US3), providing both modalities simultaneously produced efficiency gains that substantially exceeded either modality alone: correct diagnoses per minute increased by 40% and comment editing time fell by 67%, without compromising diagnostic accuracy. Notably, neither modality improved efficiency during guidance in US2, which makes the in-guidance efficiency gain under combined guidance in US3 the more striking result. Expert-distilled multimodal guidance can remove two distinct clinical workflow bottlenecks at once (visual search overhead and documentation burden) without compromising the diagnostic accuracy clinicians already achieve.
Comments23 pages, 11 figures. To appear in UIST '26: Proceedings of the 39th Annual ACM Symposium on User Interface Software and Technology, November 02-05, 2026, Detroit, MI, USA. DOI: 10.1145/3830398.3830722