arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PanDent:面向牙科放射学中全面的牙级结构-语言一致性

PanDent: Toward Comprehensive Tooth-Level Structure-Language Consistency in Dental Radiology

Xiaohan Li, Xinyu Liu, Chang Liu, Sum Wing Au Yeung, Jun Liu, Yixuan Yuan, Hui Chen

arXiv 2607.27378首次发表:更新:

发表机构

Faculty of Dentistry, The University of Hong Kong; Imperial College London; University of Science and Technology of China; Department of Data and Systems Engineering, The University of Hong Kong; Department of Electronic Engineering, The Chinese University of Hong Kong(香港大学牙医学院; 帝国理工学院; 中国科学技术大学; 香港大学数据与系统工程系; 香港中文大学电子工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出PanDent牙科OPG基准,经实验发现现有MLLM生成的牙科报告流畅但临床一致性差,在PanDent上微调可提升其结构-语言一致性,该基准可用于评估MLLM的牙级临床推理能力。

AI 中文摘要

对牙科全景X线片(全口曲面断层片,OPG)的多模态大语言模型(MLLM)进行准确评估,因缺乏反映专家解读的细粒度、临床可靠基准而受到限制。本研究推出PanDent,这是一个基于细粒度、经专家验证的牙级标注构建的大规模、临床导向的OPG基准。该数据集包含9524张高质量OPG,每张都配有由经验丰富的牙医生成、并经口腔颌面放射科医生进一步验证的全面结构化标注,为牙级诊断与推理提供临床可靠的监督。依据临床医生定义的报告逻辑,从经专家验证的发现中构建出临床一致的放射学报告,在结构化临床证据与自由文本描述之间建立明确对应关系。此设计可用于评估MLLM生成的报告是否不仅在语言上连贯,且在临床上与经专家验证的牙级发现一致。实验在多种MLLM上开展,包括最先进(SOTA)的专有模型、通用领域开源模型及医疗专用模型。结果显示,当前MLLM可生成流畅的报告,但无法产生临床一致的描述,在细粒度定位与牙级诊断中存在大量错误。在PanDent上进行微调可显著提升结构-语言一致性,大幅提高视觉定位准确率与诊断正确性,使模型输出更接近专家牙科解读。这些结果确立PanDent为评估MLLM牙级临床推理的严谨基准,以及临床导向牙科AI的宝贵资源。

英文摘要

Accurate evaluation of multimodal large language models (MLLMs) in dental panoramic radiography (orthopantomogram, OPG) is limited by the lack of fine-grained, clinically reliable benchmarks that reflect expert interpretation. This work introduces PanDent, a large-scale, clinically grounded OPG benchmark built upon fine-grained, expert-validated tooth-level annotations. The dataset comprises 9,524 high-quality OPGs, each associated with comprehensive structured annotations produced by experienced dentists and further validated by an oral and maxillofacial radiologist, providing clinically reliable supervision for tooth-level diagnosis and reasoning. Clinically consistent radiology reports are constructed from expert-validated findings using clinician-defined reporting logic, establishing explicit correspondence between structured clinical evidence and free-text descriptions. This design enables evaluation of whether MLLMs generate reports that are not only linguistically coherent but also clinically consistent with expert-validated tooth-level findings. Experiments are conducted on diverse MLLMs, including state-of-the-art (SOTA) proprietary models, general-domain open-source models, and medical-specific models. Results show that current MLLMs can generate fluent reports, yet fail to produce clinically consistent descriptions, exhibiting substantial errors in fine-grained localization and tooth-level diagnosis. Fine-tuning on PanDent significantly improves structure-language consistency, substantially enhancing visual localization accuracy and diagnostic correctness, and bringing model outputs closer to expert dental interpretation. These results establish PanDent as a rigorous benchmark for evaluating tooth-level clinical reasoning in MLLMs and a valuable resource for clinically grounded dental AI.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑