arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

在翻转之前:量化视觉语言模型中答案改变前的隐藏分数偏移测量用于视觉问答

BEFORE THE FLIP: Measuring Hidden Score Shifts In Quantized Vision Language Models Before The Answer Changes for Visual Question Answering

Sourajit Saha, Shubhashis Roy Dipta, Shaswati Saha, Nobin Sarwar, Yuxuan Jiang

arXiv 2609.06922首次发表:更新:

发表机构

University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出“在翻转之前”方法,测量量化视觉语言模型中答案不变时的隐藏分数偏移,发现4比特压缩比8比特更显著改变分数差距,但逐问题调整精度无可靠益处。

AI 中文摘要

量化通过使用更少的比特来表示视觉语言模型(VLM)的权重,使其存储和运行更加廉价。虽然压缩后视觉问答(VQA)的答案保持不变是预期行为,但这仍可能隐藏底层分数(对数概率)的变化。例如,模型在压缩后可能仍然回答“是”,即使“是”和“否”之间的分数差距已经缩小。我们引入了“在翻转之前”方法来测量这些隐藏变化。我们的方法将压缩引起的分数变化与用固定的平均令牌替换图像内部表示(即图像令牌)引起的变化进行比较。然后,我们一次增加一个权重组的精度,以确定额外比特在何处有帮助,并测试为每个问题选择不同组是否比打乱的控制组更有优势。在8,277个LLaVA问题中,图像令牌替换可测量地影响分数,4比特压缩将“是”或“否”分数差距推向替换输出的程度比8比特压缩更远。Qwen显示出相同的模式,但差异较小。然而,在9,000个LLaVA答案中,只有265个在4比特下发生变化。在一项针对1,024个校准问题的单独研究中,为每个问题单独选择权重组在任何测试的存储预算下均未优于两个打乱的控制组。这些发现表明,压缩可以改变不变答案背后的分数,但并未确立为每个问题调整精度的可靠益处。

英文摘要

Quantization makes vision language models (VLMs) cheaper to store and run by using fewer bits to represent their weights. While unchanged answers on visual question answering (VQA) after compression are an expected behavior, they can still hide changes in the underlying scores (log probabilities). For example, a model may still answer yes after compression, even as the score gap between yes and no shrinks. We introduce BEFORE THE FLIP to measure these hidden changes. Our method compares the score change caused by compression with the change caused by replacing the image's internal representations, or image tokens, with one fixed average token. We then increase the precision of one weight group at a time to identify where extra bits help, and test whether choosing different groups for each question offers benefits beyond shuffled controls. Among 8,277 LLaVA questions where image token replacement measurably affects the scores, 4-bit compression shifts the yes or no score gap farther toward the replacement output than 8-bit compression. Qwen shows the same pattern, but with a smaller difference. Yet only 265 of 9,000 LLaVA answers change at 4 bits. In a separate study of 1,024 calibration questions, choosing weight groups separately for each question does not outperform both shuffled controls at any tested storage budget. These findings show that compression can alter the scores behind unchanged answers, but do not establish a reliable benefit from adjusting precision for each question.

CommentsUnder Review at VLM4RWD @ NeurIPS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑