AI 中文总结
该研究提出WEP方法,利用MLLM的白盒预填充路径,无需额外操作即可检测闭集视觉答案的推理风险,在3个MLLM和4个基准上提升了平均错误AP。
AI 中文摘要
多模态大语言模型(MLLM)的可靠部署需要判断一个有把握的视觉答案是否可信、是否需要审核,或是是否应路由至更强的系统。置信分数捕捉候选答案的边际,但不捕捉与这些边际相关的估计符号化视觉读数的来源或分布。我们利用生成答案的相同白盒预填充路径,研究闭集视觉答案的推理时风险检测。证据见证组合(WEP)逐层估计支持或矛盾于预测候选答案的视觉贡献,通过两类可解释路由族汇总这些贡献:与问题相关的证据来源,以及符号化证据集中度。嵌套分组验证选择更可靠的路由族和稀疏的前k个路由组合,将其与候选置信分数融合。WEP无需图像扰动、解码变更、反向传播或外部验证器。在3个MLLM和4个二值答案基准上,WEP将平均错误AP提升0.134;全部12组模型-数据集组合的增益均为正,且10对的图像聚类自举区间严格为正。WEP针对白盒闭集答案系统,使用带标签的校准切片。
英文摘要
Reliable deployment of multimodal large language models (MLLMs) requires deciding whether a confident visual answer should be trusted, reviewed, or routed to a stronger system. Confidence scores capture candidate margins, but not where the estimated signed visual readouts associated with those margins come from or how they are distributed. We study inference-time risk detection for closed visual answers using the same white-box prefill path that produces the answer. Witness Evidence Portfolios (WEP) first estimates, layer by layer, which visual contributions support or contradict the predicted candidate. It summarizes these contributions through two interpretable route families: question-related evidence provenance and signed evidence concentration. Nested grouped validation chooses the more reliable family and a sparse top-k route portfolio, which is fused with candidate confidence. WEP needs no image perturbation, decoding change, backward pass, or external verifier. Across three MLLMs and four binary-answer benchmarks, WEP improves mean error AP by 0.134. All 12 model--dataset gains are positive, and image-cluster bootstrap intervals are strictly positive on 10 pairs. WEP targets white-box closed-answer systems and uses a labeled calibration slice.
Comments22 pages, 6 figures; includes supplementary material. Code: https://github.com/SouthWinter/WEP