AI 中文总结
研究旨在解决现有医学视觉问答基准测试无法充分体现全景牙科X光片解读复杂性的问题,引入DentiAsk基准测试,涵盖多种病变和推理层次,对10种模型测试发现其在空间定位等方面存在局限,为推进医学成像多模态推理提供挑战基准。
AI 中文摘要
准确解读全景牙科X光片需要整合多种推理能力,包括检测、空间定位和定量评估。尽管多模态学习取得了进展,但现有的医学视觉问答基准测试未能充分体现这种复杂性。为此,我们引入了DentiAsk,这是一个大规模牙科视觉问答基准测试,它将高分辨率全景牙科X光片与经过临床医生验证的问答对配对,涵盖三个推理层次和三种高发性病变。DentiAsk包含1000张高分辨率X光片及10000个专家策划的问答对。我们对10种先进视觉语言模型进行基准测试,发现模型在描述性查询上表现较好,但在空间定位和计数方面表现大幅下降,揭示了视觉识别与临床有意义推理之间的差距,确立了DentiAsk作为推进医学成像中多模态推理的具有挑战性的基准测试。
英文摘要
Accurate interpretation of panoramic dental radiographs requires the integration of multiple reasoning capabilities: detection, spatial localization, and quantitative assessment. Despite recent advances in multimodal learning, existing medical visual question answering (VQA) benchmarks do not fully capture this complexity, often reducing the task to simplified classification or templated queries. As a result, they provide limited coverage of the diverse reasoning processes required for clinically meaningful interpretation. We introduce DentiAsk, a large-scale dental VQA benchmark that pairs high-resolution panoramic dental radiographs with clinician-validated question-answer pairs spanning three reasoning tiers: descriptive recognition, spatial localization, and numerical quantification across three high-prevalence pathologies: periapical radiolucency (PARL), impacted teeth, and dental caries. DentiAsk comprises 1,000 high-resolution radiographs annotated with 10,000 expert-curated QA pairs. To our knowledge, it is the first dental VQA benchmark to unify categorical, spatial, and quantitative reasoning as separately scored tasks within a single evaluation framework. We benchmark 10 state-of-the-art vision-language models, including LLaVA-v1.5, LLaVA-v1.6, Qwen-VL, InternVL2, and LLaVA-Med, and find that models achieve stronger performance on descriptive queries, whereas they degrade sharply on spatial localization and counting, exposing limitations in compositional, multi-step reasoning. These findings reveal a gap between visual recognition and clinically meaningful reasoning, establishing DentiAsk as a challenging benchmark for advancing multimodal reasoning in medical imaging.