arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17800cs.CV

AgenTeeth:一种通过工具证据注入抑制冻结视觉语言模型在牙科X光片上幻觉的模型无关框架

AgenTeeth: A Model-Agnostic Framework for Suppressing Hallucination in Frozen Vision-Language Models on Dental X-Rays via Tool Evidence Injection

Ahmed Rafid, Fariya Ahmed, Rumman Adib, Mehedi Ahamed, Ajwad Abrar, Tareque Mohmud Chowdhury

首次发表
浏览论文内容

中文总结 AI 辅助

AgenTeeth通过七个牙科视觉专家工具注入证据,无需微调即可抑制冻结VLM在牙科X光片上的幻觉,在MMOral-OPG-Bench上提升12.9-23.0个百分点,并超越OralGPT-Plus。

中文摘要 AI 辅助

视觉语言模型(VLM)在全景牙科X光片上仍然在很大程度上不可靠,可能依赖学习到的解剖学先验而非图像中的证据。这在牙齿定位和空间推理方面尤其成问题,并且经过微调的牙科VLM可能保留相同的空间偏差。我们提出了AgenTeeth,一个模型无关的、工具增强的框架,利用七个专门的牙科视觉专家来锚定冻结的VLM。一个问题感知的编排器选择相关工具,其检测结果映射到FDI牙齿编号或解剖区域,并连同带注释的图像叠加层作为结构化发现返回。然后,一次新的综合调用利用这些证据回答问题,而无需微调底层VLM。在MMOral-OPG-Bench上,AgenTeeth将四个骨干VLM相对于其基线提高了12.9-23.0个百分点。我们最强的配置在开放式VQA上达到65.66%,而OralGPT-Plus为45.35%。在匹配规模下优势同样存在:冻结的Qwen2.5-VL-7B-Instruct与AgenTeeth结合达到48.11%,超过了基于同一骨干经过监督微调和强化学习用于工具使用的OralGPT-Plus。我们发布了该框架、全部七个专家模型,以及一个由牙医标注的用于全景X光片中牙槽骨丢失检测的数据集。

英文摘要

Vision-language models (VLMs) remain largely unreliable on panoramic dental radiographs and can rely on learned anatomical priors rather than evidence in the image. This is particularly problematic for tooth localization and spatial reasoning, and fine-tuned dental VLMs can retain the same spatial biases. We present AgenTeeth, a model-agnostic, tool-augmented framework that grounds frozen VLMs using seven specialized dental vision experts. A question-aware orchestrator selects the relevant tools, whose detections are mapped to FDI tooth numbers or anatomical regions and returned as structured findings together with annotated image overlays. A fresh synthesis call then answers the question using this evidence, without fine-tuning the underlying VLM. On MMOral-OPG-Bench, AgenTeeth improves four backbone VLMs by 12.9-23.0 percentage points over their baselines. Our strongest configuration reaches 65.66% on open-ended VQA, compared with 45.35% for OralGPT-Plus. The advantage also holds at matched scale: a frozen Qwen2.5-VL-7B-Instruct with AgenTeeth reaches 48.11%, exceeding OralGPT-Plus built on the same backbone after supervised fine-tuning and reinforcement learning for tool use. We release the framework, all seven expert models, and a dentist-annotated dataset for alveolar bone-loss detection in panoramic radiographs.

发表机构

  • Islamic University of Technology(伊斯兰理工大学)
  • Southeast University(东南大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑