arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型能否解锁眼科诊断报告中的离散数据?

Can large language models unlock discrete data in ophthalmic diagnostic reports?

Umair A. Zaidi, An-Lun Wu, Wei-Chun Lin, Thomas S. Hwang, Michelle R. Hribar

arXiv 2610.00795首次发表:更新:

发表机构

Oregon Health & Science University; Casey Eye Institute; University of Illinois Chicago(俄勒冈健康与科学大学; 凯西眼科研究所; 伊利诺伊大学芝加哥分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估GPT-4o辅助流水线从眼科诊断PDF中提取结构化数据,发现两种策略均高准确且耗时减少约92%,Prompt-Only值准确性最高,Schema-Constrained格式完全合规。

AI 中文摘要

目的:评估采用两种提示策略的大型语言模型(LLM)从眼科诊断PDF报告中提取结构化数据的准确性和效率。方法:使用两条GPT-4o辅助流水线处理四种类型(视野、OCT青光眼概览、OCT视网膜神经纤维层单次检查、OCT厚度图;各5份)的20份去标识化报告,并与经过核对的人工金标准进行比较。Schema-Constrained使用结构化输出模式并预定义JSON Schema;Prompt-Only使用详细指令提示,随后通过Python转换为JSON。评估结果为值准确性、格式准确性和提取时间。结果:Schema-Constrained在视野和RNFL单次检查中的值准确性为100.00%,青光眼概览为97.45%,厚度图为98.00%;Prompt-Only在所有四种报告类型中均达到100.00%。格式准确性方面,Schema-Constrained在所有报告类型中均为100.00%,Prompt-Only除RNFL单次检查(90.14%)外均为100.00%。人工复核的平均提取时间为每份报告56.51秒,而Schema-Constrained为5.04秒,Prompt-Only为4.70秒,约减少92%。结论:在这个小型概念验证数据集中,通用LLM辅助流水线从眼科诊断PDF中提取结构化数据具有高准确率,并大幅缩短处理时间。Prompt-Only达到最高值准确性,而Schema-Constrained产生符合Schema的输出,格式准确性为100%。这些互补优势支持进一步评估用于研究和临床数据提取的混合式、带验证的工作流。

英文摘要

Objective: To assess the accuracy and efficiency of a large language model (LLM) using two prompt strategies to extract structured data from ophthalmic diagnostic PDF reports. Methods: Twenty deidentified reports across four types (Visual Field, OCT Glaucoma Overview, OCT retinal nerve fiber layer Single Exam, and OCT Thickness Map; n = 5 each) were processed using two GPT-4o-assisted pipelines and compared with a reconciled manual ground truth. Schema-Constrained used Structured Output mode with a predefined JSON Schema; Prompt-Only used a detailed instruction prompt followed by Python conversion to JSON. Outcomes were value accuracy, formatting accuracy, and extraction time. Results: Schema-Constrained value accuracy was 100.00% for Visual Field and RNFL Single Exam, 97.45% for Glaucoma Overview, and 98.00% for Thickness Map; Prompt-Only achieved 100.00% across all four report types. Formatting accuracy was 100.00% for Schema-Constrained across all report types and 100.00% for Prompt-Only except RNFL Single Exam (90.14%). Mean extraction time was 56.51 s per report for manual review versus 5.04 s for Schema-Constrained and 4.70 s for Prompt-Only, an approximately 92% reduction. Conclusions: In this small proof-of-concept dataset, general-purpose LLM-assisted pipelines extracted structured data from ophthalmic diagnostic PDFs with high accuracy and substantially reduced processing time. Prompt-Only achieved the highest value accuracy, while Schema-Constrained produced schema-compliant output with 100% formatting accuracy. These complementary strengths support further evaluation of hybrid, validation-aware workflows for research and clinical data abstraction.

Comments10 pages, 5 figures. Presented at the Association for Research in Vision and Ophthalmology (ARVO) Annual Meeting, Denver, Colorado, May 4, 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑