AI 中文总结
研究内窥镜VQA中复杂问题答案一致性,引入EndoCA基准,评估11个VLM,发现部分模型复杂答案准确率高但原子答案准确率及复杂-原子答案一致性低,提出ASR机制,能在选定模型上提高配对复杂-原子正确性及已回答案例准确性。
AI 中文摘要
内窥镜视觉问答(VQA)越来越多地提出结合多个内窥镜答案成分的复杂问题,而非孤立的事实性查询。即使同一模型在相关原子问题上失败,此类复杂答案也可能被判定为正确。我们引入EndoCA,一个用于评估复杂答案与同图像原子答案是否一致的配对复杂-原子答案一致性基准。EndoCA包含两个套件:EndoCA-Core评估实际内窥镜VQA中常见的紧凑问题-复杂度模式,EndoCA-Diagnostic支持跨增加问题复杂度的受控分析。我们在EndoCA上评估了11个VLM。为减少复杂-原子不一致性,我们引入原子支持协调(ASR),一种无需训练的机制。在四个选定的公开可用模型上,ASR-Revise以适度改变复杂答案准确性的方式提高配对复杂-原子正确性,而ASR-Selective通过允许模型放弃不太可靠的情况提高已回答案例的准确性。总之,EndoCA和ASR为内窥镜VQA中的答案协调和选择性回答提供了一个一致性感知基准和一种无需训练的机制。
英文摘要
Endoscopic visual question answering (VQA) increasingly asks complex questions that combine several endoscopic answer components rather than isolated factual queries. Such complex answers may be scored as correct even when the same model fails on associated atomic questions. We introduce EndoCA, a paired complex-atomic answer consistency benchmark for evaluating whether complex answers remain consistent with same-image atomic answers. EndoCA contains two suites: EndoCA-Core evaluates compact question-complexity patterns commonly seen in practical endoscopic VQA, and EndoCA-Diagnostic supports controlled analysis across increasing question complexity. We evaluate 11 VLMs spanning open, medical, endoscopy-adapted, and closed-source models on EndoCA. Some VLMs achieve high complex-answer accuracy, yet their atomic-answer accuracy and complex-atomic answer consistency remain substantially lower. To reduce this complex-atomic inconsistency, we introduce Atomic-Support Reconciliation (ASR), a training-free mechanism that uses model-generated atomic answers as contextual premises for answer revision and consistency-guided selective answering. On four selected publicly available models, ASR-Revise improves paired complex-atomic correctness with modest changes in complex-answer accuracy, while ASR-Selective improves accuracy on answered cases by allowing the model to abstain from less reliable cases. Together, EndoCA and ASR provide a consistency-aware benchmark and a training-free mechanism for answer reconciliation and selective answering in endoscopic VQA.
Comments7 pages, 5 figures, 6 tables