arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.00077cs.CV

超越准确率:针对OCR关键多模态大语言模型(MLLM)推理的视觉令牌剪枝空间溯源审计

Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference

Feixiang Liu, Qiang Qiu, Hao Zhang, Xinyue Wang

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对OCR关键MLLM推理的视觉令牌剪枝,提出结合空间溯源的审计方法,发现准确率无法反映的质量风险,揭示模型特定前沿并验证压缩与任务通用性的关系,为视觉令牌剪枝评估提供新范式。

中文摘要 AI 辅助

视觉令牌剪枝通常以固定保留预算下的答案质量作为评判标准。对于富含文本的多模态大语言模型(MLLM),该评判协议会忽略一种独特的失效情况:即使答案正确,也不存在任何保留的令牌可在局部追溯到支撑该答案的小型OCR区域。我们将这一盲区转化为证据风险审计,该审计将答案行为与几何令牌起源溯源、干预措施及实际成本相结合;透明的无训练选择器可分离出可控操作点。在锁定图像不相交确认任务中,30%保留率的Qwen Target模型观测到的准确率为0.786,而Full模型为0.783(配对图像集群差异为+0.003,95%置信区间为[-0.014, +0.020]),但相同预算下的Target、Random及Grid模型的正支持覆盖率差异显著,分别为0.620、0.270和0.318。在Qwen3-VL-8B、LLaVA-1.5-7B及InternVL3.5-8B模型上,匹配对照、干预措施、检测器测试及外部方法揭示了仅靠准确率无法暴露的模型特定质量风险溯源前沿。物化前缀可实现最高4.32倍的批次预填充加速及76.4%的增量峰值内存降低;全验证的TextVQA和DocVQA进一步表明,有利的目标验证点并不意味着任务通用压缩。因此,视觉令牌剪枝应在报告质量和压缩率的同时,报告存活的空间溯源及实际成本。

英文摘要

Visual-token pruning is usually judged by answer quality at a fixed retention budget. For text-rich multimodal large language models (MLLMs), this protocol can miss a distinct failure: an answer remains correct even when no retained token is locally traceable to the small OCR region that supports it. We turn this blind spot into an evidence-risk audit that couples answer behavior with geometric token-origin provenance, interventions, and realized cost; transparent training-free selectors isolate controlled operating points. On locked image-disjoint confirmation, Qwen Target at 30% retention has observed accuracy 0.786 versus 0.783 for Full (paired image-cluster difference +0.003, 95% CI [-0.014, +0.020]), yet same-budget Target, Random, and Grid retain sharply different positive-support coverage: 0.620, 0.270, and 0.318. Across Qwen3-VL-8B, LLaVA-1.5-7B, and InternVL3.5-8B, matched controls, interventions, detector tests, and external methods reveal model-specific quality-risk-traceability frontiers that accuracy alone does not expose. Materialized prefixes yield up to 4.32x batch-prefill speedup and 76.4% lower incremental peak memory; full-validation TextVQA and DocVQA further show that favorable target-verification points do not imply task-general compression. Visual-token pruning should therefore report surviving spatial provenance and realized cost alongside quality and compression.

补充信息

↑