AI 中文总结
针对视觉-语言模型的常识驱动幻觉,提出选择性先验校准方法,可提升反事实图像准确率并保留常识图像准确率,增益可推广至多类基准。
AI 中文摘要
在视觉-语言模型中,常识驱动的幻觉(CDH)会在模型的常识先验覆盖异常状态的清晰视觉证据时发生,例如模型可能会将明显有六根手指的手报告为五根。我们表明这些错误具有系统性:当模型对反事实(CF)图像的问题回答错误时,其答案往往与不访问图像时偏好的候选答案一致。不加区分地抑制该先验可修复CF错误,但也可能破坏相同先验有帮助的匹配常识(CS)图像的正确答案。因此我们提出选择性先验校准(SPC),该方法以实例依赖的强度从图像条件得分中减去候选级先验偏好估计,且仅当所得得分模式强烈支持替代方案时才修改原始预测。大量实验表明,SPC大幅提高了CF图像的准确率,同时基本保留了匹配CS图像的准确率;此外,这些增益可推广至CDH类别、候选答案排列及其他冲突基准,且SPC极少改变无此类冲突基准的预测。
英文摘要
In vision--language models, commonsense-driven hallucination (CDH) occurs when a model's commonsense prior overrides clear visual evidence of an atypical state. For example, a model may report that a visibly six-fingered hand has five fingers. We show that these errors are systematically directed: when a model answers a question about a counterfactual (CF) image incorrectly, its answer often coincides with the candidate it prefers without access to the image. Suppressing this prior indiscriminately can repair CF errors, but may also disrupt correct answers on matched commonsense (CS) images, where the same prior is helpful. We therefore propose Selective Prior Calibration (SPC), which subtracts candidate-level prior-preference estimates from image-conditioned scores with an instance-dependent strength and revises the original prediction only when the resulting score pattern strongly supports an alternative. Extensive experiments demonstrate that SPC substantially improves accuracy on CF images while largely preserving accuracy on matched CS images. Furthermore, these gains generalize across CDH categories, candidate-answer permutations, and other conflict benchmarks, while SPC rarely alters predictions on benchmarks without such conflicts.