arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更大还是更便宜?图像退化下视觉语言模型中尺度和量化对不确定性信号的影响

Bigger or Cheaper? Scale and Quantization Effects on Uncertainty Signals in Vision-Language Models Under Image Degradation

M M Asif Ferdous

arXiv 2607.24440首次发表:更新:

AI 中文总结

研究在图像退化下视觉语言模型中尺度和量化对不确定性信号的影响,通过对Qwen2-VL系列模型的实验发现,尺度提升内部不确定性信号,4位量化对准确性影响小但对置信信号影响大,建议固定内存预算下选更大量化模型,如7B-4bit。

AI 中文摘要

部署在消费硬件上的视觉语言模型(VLM)必须决定何时回答和何时延迟,而该决定取决于具有跟踪正确性的置信信号。有固定内存预算的从业者面临着三种配置的选择:全精度的小模型、量化后的相同小模型,以及量化后占用相同空间的更大模型,这三种配置会使置信信号朝相反方向变化。我们在相同输入上测量模型尺度和4位量化如何影响Qwen2-VL系列中的两个置信信号:模型在自然语言中陈述的置信度,以及其在生成答案上的平均token概率。在跨越三种严重程度的六种现实照片退化的5700个预测中,我们发现尺度大幅改善了模型的内部不确定性信号(平均错误检测AUROC从2B模型的0.80提升到7B模型的0.98),而其语言表达的置信度仍然较弱且通常处于随机水平(平均从0.61到0.69):模型所知与所说之间的差距随着模型大小的增加而扩大而非缩小。我们发现4位量化对准确性几乎没有影响(下降1.6个点),但对置信信号影响很大(内部AUROC从0.95降至0.80,语言表达置信度的解析率从99%降至64%)。因此,对于固定内存预算,建议选择更大的量化模型而非更小的全精度模型:7B-4bit在三种适用配置中具有最佳准确性和最佳不确定性信号(内部AUROC为0.98)。我们将结果构建为选择性预测操作点,以便直接转化为部署建议,并认为错误检测AUROC而非校准误差是揭示两个信号差异的指标。

英文摘要

Vision-language models (VLMs) deployed on consumer hardware must decide when to answer and when to defer, and that decision depends on having a confidence signal that tracks correctness. A practitioner with a fixed memory budget faces a choice between a small model at full precision, the same small model quantized, and a larger model quantized into the same footprint -- three configurations that push the confidence signal in opposing directions. We measure, on identical inputs, how model scale and 4-bit quantization affect two confidence signals in the Qwen2-VL family: the confidence a model states in natural language, and its own mean token probability over the answer it generates. Across 5,700 predictions spanning six realistic photographic degradations at three severities, we find that scale sharply improves the model's internal uncertainty signal (mean error-detection AUROC 0.80 to 0.98 from 2B to 7B) while its verbalized confidence stays weak and often at chance (mean 0.61 to 0.69): the gap between what the model knows and what it says widens rather than closes with size. We find that 4-bit quantization is nearly free for accuracy (-1.6 points) but expensive for the confidence signal (internal AUROC 0.95 to 0.80, and the verbalized-confidence parse rate collapses from 99% to 64%). For a fixed memory budget the recommendation is therefore to prefer a larger quantized model over a smaller full-precision one: 7B-4bit gives both the best accuracy and the best uncertainty signal (internal AUROC 0.98) of the three configurations that fit. We frame the results as selective-prediction operating points so they translate directly into a deployment recommendation, and we argue that error-detection AUROC, not calibration error, is the metric that exposes the difference between the two signals.

Comments12 pages, 4 figures. Code and data: https://github.com/Asif-Ferdous/vlm-scale-quant

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑