arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29999cs.CVcs.AIcs.LGcs.MM

GHOST-Q:探究量化视觉语言模型在同分权衡下被忽视的接地幻觉

GHOST-Q: Towards Studying Grounding Hallucinations Overlooked Under Same-score TradeOffs in Quantized VLMS

发表机构纽约大学阿布扎比分校
查看机构详情
  • New York University Abu Dhabi (NYUAD)(纽约大学阿布扎比分校)

机构由 AI 辅助整理,请以论文原文为准。

Saim Rehman, Muhammad Shafique

首次发表
浏览论文内容

中文总结 AI 辅助

GHOST-Q通过跨精度配对评估,揭示量化VLM在保持总体准确率的同时,接地幻觉显著增加,且内存减少不等于延迟降低,需综合评估效用、接地、生成与部署效率。

中文摘要 AI 辅助

视觉语言模型(VLM)的后训练量化通常通过总体任务准确率和内存节省来评估,但保持一个亮眼的分数并不能保证视觉接地行为的保持。我们提出了GHOST-Q,一种跨精度受控评估方法,对三个8B VLM系列在FP16、INT8和NF4精度下,在效用和幻觉敏感基准上进行评估。我们不仅比较总体准确率,还逐项配对FP16和量化预测,以量化压缩如何重新分配接地成功与失败。六个量化变体中有五个在MMStar准确率上保持在±2个百分点以内,但在错误发现率校正后,36个配对效应中有10个仍然显著,其中九个出现在幻觉敏感条件下。同设备A100性能分析进一步表明,显著的内存减少并不一定意味着更低的推理延迟。最后,一项开放式AMBER审计揭示了强烈的生成预算审查,其严重程度因架构和精度而异。这些结果表明,量化VLM应综合评估其总体效用、接地可靠性、生成行为和实际部署效率。

英文摘要

Post-training quantization of vision--language models (VLMs) is typically assessed through aggregate task accuracy and memory savings, but preserving a headline score does not guarantee preservation of visual grounding behavior. We present GHOST-Q, a cross-precision controlled evaluation of three 8B VLM families under FP16, INT8, and NF4 across utility and hallucination-sensitive benchmarks. Rather than comparing only aggregate accuracy, we pair FP16 and quantized predictions item by-item to quantify how compression redistributes grounding successes and failures. Five of six quantized variants preserve MMStar accuracy within $\pm2$ percentage points, yet 10 of 36 paired effects remain significant after false-discovery-rate correction, nine on hallucination-sensitive conditions. Same-device A100 profiling further demonstrates that substantial memory reduction does not necessarily mean lower inference latency. Finally, an open-ended AMBER audit reveals strong generation budget censoring whose severity varies by architecture and precision. These results show that quantized VLMs should be evaluated jointly for aggregate utility, grounding reliability, generation behavior, and realized deployment efficiency.

补充信息

↑