EVADE:带弃权机制的证据验证智能诊断方法
EVADE: Evidence-Verified Agentic Diagnosis with Escape
- Monash University(莫纳什大学)
- Murdoch University(默多克大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对医学VLMs不可靠问题,提出无需训练的EVADE方法,通过跨图像视图验证一致性并引入弃权机制,在多医学VQA数据集上提升了校准度与选择性风险,同时维持了准确率。
AI中文摘要:
医学视觉语言模型(VLMs)可达到较高准确率,但仍存在不可靠性:它们系统性地过度自信,几乎无法从测试时推理中获益,且缺乏可靠校准自身响应可信度的能力。我们提出EVADE(带弃权机制的证据验证智能诊断方法),这是一种无需训练的推理式方法,可提升单个冻结VLMs部署的安全性。EVADE先给出响应,当存在不确定性时,定位最具诊断相关性的区域,在放大视图上重新作答,且仅在整图与放大视图的响应一致时才确定答案;若不一致则弃权(不执行)。为直接解决单模型自检查中的验证幻觉问题,我们的核心思路是跨不同图像视图验证门控一致性,而非重读模型自身文本。在VQA-RAD、SLAKE和PathVQA数据集上使用Qwen2.5-VL-7B开展的实验评估显示,EVADE是唯一同时提升校准度与选择性风险、并维持准确率的方法,与零样本方法相比,其预期校准误差(ECE)最多降低45%。思维链、自一致性与自验证方法均至少在一个维度上失效。定位分析表明,自提议区域在诊断结构定位上优于中心区域或随机裁剪区域,但7B规模的VLM无法利用该定位结果修正答案。因此,可靠性提升源于一致性门控与经校准的弃权(不执行)机制。
英文摘要:
Medical vision-language models (VLMs) can achieve high accuracy but remain unreliable: they are systematically overconfident, benefit little from test-time reasoning, and lack the ability to reliably calibrate trust in their own responses. We introduce EVADE (Evidence-Verified Agentic Diagnosis with Escape), an inferential, non-training method that enhances the safety of deploying a single frozen VLM. EVADE responds and, when uncertain, localises the region most diagnostically relevant, re-answers on a zoomed view, and commits only when both the entire image and the zoomed view responses agree; otherwise, it abstains. To directly address verification hallucination in single-model self-checking, our main idea is to verify gate consistency across different image views rather than re-reading the model's own text. Experimental evaluation on VQA-RAD, SLAKE, and PathVQA using Qwen2.5-VL-7B reports that EVADE is the only method that simultaneously improves both calibration and selective risk while maintaining accuracy, reducing expected calibration error (ECE) by up to 45% compared to zero-shot. Chain-of-thought, self-consistency, and self-verification all fail at least one axis. A grounding analysis reports that self-proposed regions perform better at diagnostic structure localisation than centres or random crops. However, a 7B VLM cannot use this localisation to revise answers. Therefore, reliability gains come from the consistency gate and calibrated abstention.