AI 中文总结
针对VLM部署中预测信任问题,提出模型无关的ATS框架,聚合8种诊断分数评估答案可靠性,可识别稳定错判等失效模式,补充模型级能力评估。
AI 中文摘要
视觉-语言模型(VLMs)可从视觉观测中估计时长、速度、加速度等物理量,但现有基准主要评估模型针对标注真实值的整体性能。部署时的关键问题是,当真实值不可用时,能否信任单个预测结果。仅靠自洽性可能无法捕捉重要失效模式:VLM可能生成稳定但错误的估计,或依赖文本先验而非视觉证据。我们针对定量物理推理提出答案级选择性预测,以及Answer-Level Trust Selection(ATS)——一种事后、模型无关的框架,用于接受或拒绝单个VLM预测。ATS无需微调、辅助验证器或访问模型内部logits,而是将从重复查询和受控干预中得到的8个可解释行为诊断分数聚合为统一信任分数。我们在Qwen2.5-VL-7B及20种VLM主干上对ATS进行深度评估,考察其选择性性能、诊断行为和针对性失效模式。结果显示,基于干预的诊断有助于识别仅靠重复一致性可能遗漏的稳定但错误预测和先验追踪预测;不过,对失效案例的拒绝性能提升可能伴随正确预测保留率降低的代价。因此,ATS通过答案级可靠性评估,补充了针对定量VLM预测的模型级能力评估,代码将在发表后发布。
英文摘要
Vision-language models (VLMs) can estimate physical quantities such as duration, speed, and acceleration from visual observations, but existing benchmarks primarily assess overall model performance against annotated ground truth. In deployment, a key question is whether an individual prediction can be trusted when its ground truth is unavailable. Self-consistency alone may fail to capture important failure modes: a VLM may produce stable-but-wrong estimates or rely on textual priors rather than visual evidence. We formulate answer-level selective prediction for quantitative physical reasoning and propose Answer-Level Trust Selection (ATS), a post-hoc, model-agnostic framework for accepting or rejecting individual VLM predictions. ATS requires no fine-tuning, auxiliary verifier, or access to the model's internal logits. Instead, it aggregates eight interpretable behavioral diagnostic scores derived from repeated queries and controlled interventions into a unified trust score. We evaluate ATS in depth on Qwen2.5-VL-7B and across 20 VLM backbones, examining selective performance, diagnostic behavior, and targeted failure modes. Our results show that intervention-based diagnostics help identify stable-but-wrong and prior-tracking predictions that repeated agreement alone may miss. However, improved failure-case rejection can come at the cost of lower retention of correct predictions. ATS therefore complements model-level capability evaluation with answer-level reliability assessment for quantitative VLM predictions. Code will be released upon publication.
CommentsPreprint