arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.34044cs.CV

SCOPD:面向高效视觉语言模型的稀疏上下文在线策略自蒸馏

SCOPD: Sparse-Context On-Policy Self-Distillation for Efficient Vision-Language Models

  • University of Toronto(多伦多大学)
  • Vector Institute(向量研究所)
  • Samsung AI Center Toronto(三星多伦多人工智能中心)
  • Stanford University(斯坦福大学)
  • University of Waterloo(滑铁卢大学)
  • University of British Columbia(不列颠哥伦比亚大学)
  • Alan Turing Institute(艾伦·图灵研究所)
  • Tsinghua University(清华大学)
  • York University(约克大学)
  • NVIDIA(英伟达)

机构由 AI 辅助整理,请以论文原文为准。

Ahmadreza Jeddi, Enming Zhang, Jasper Gerigk, Hakki Karaimer, Mozhgan Nasr Azadani, Jiayun Luo, Minh Ngoc Le, Gholamali Aminian, Hugo Buurmeijer, Yongchao Chen,… 展开作者

Ahmadreza Jeddi, Enming Zhang, Jasper Gerigk, Hakki Karaimer, Mozhgan Nasr Azadani, Jiayun Luo, Minh Ngoc Le, Gholamali Aminian, Hugo Buurmeijer, Yongchao Chen, Leonid Sigal, Igor Gilitschenski, Konstantinos G. Derpanis, Marco Pavone, Babak Taati

AI总结:

针对推理型视觉语言模型剪枝后性能下降问题,提出SCOPD及SCOPD+,通过在线策略自蒸馏和视觉敏感位置选择性蒸馏,在10%视觉令牌保留下将性能从86.37%提升至92.43%。

AI中文摘要:

推理型视觉语言模型(VLM)将图像和视频处理为长序列的视觉令牌,导致推理成本高昂。免训练的令牌剪枝可降低这一成本,但激进的压缩会急剧降低性能,这通常被归因于任务相关视觉信息的不可逆丢失。我们表明,这一解释并不完整。在固定上下文的Pass@K分析中,从同一剪枝后的视觉表示中重复采样可恢复贪婪解码所遗漏的许多示例,这表明有用的视觉证据可能仍然可访问,但使用不可靠。我们将此称为表示-利用差距。受此观察启发,我们提出SCOPD,一种稀疏上下文在线策略自蒸馏框架,其中学生模型从剪枝后的视觉令牌生成推理轨迹,而特权全上下文教师模型监督相同的在线策略前缀。SCOPD不需要真实答案、架构更改或额外的推理时计算。我们进一步引入SCOPD+,它使用小的视觉预算干预来识别视觉敏感响应位置并选择性地蒸馏它们。在10%视觉令牌保留率下,Vanilla模型在13个基准上保留其未剪枝性能的86.37%。SCOPD将其提升至90.49%,而SCOPD+进一步将其提升至92.43%。在不同令牌预算、基准和剪枝算子中,我们的结果表明,高效推理不仅取决于哪些视觉信息在剪枝中幸存,还取决于模型学习使用它的可靠性。

英文摘要:

Reasoning vision-language models (VLMs) process images and videos as long sequences of visual tokens, making inference expensive. Training-free token pruning reduces this cost, but aggressive compression can sharply degrade performance, often attributed to irreversible loss of task-relevant visual information. We show that this explanation is incomplete. In a fixed-context Pass@K analysis, repeated sampling from the same pruned visual representation recovers many examples missed by greedy decoding, indicating that useful visual evidence can remain accessible but be used unreliably. We call this the representation-utilization gap. Motivated by this observation, we introduce SCOPD, a sparse-context on-policy self-distillation framework in which a student generates reasoning trajectories from pruned visual tokens while a privileged full-context teacher supervises the same on-policy prefixes. SCOPD requires no ground-truth responses, architectural changes, or additional inference-time computation. We further introduce SCOPD+, which uses a small visual-budget intervention to identify visually sensitive response positions and selectively distill them. At 10% visual-token retention, the Vanilla model retains 86.37% of its unpruned performance across 13 benchmarks. SCOPD raises this to 90.49%, while SCOPD+ further improves it to 92.43%. Across token budgets, benchmarks, and pruning operators, our results show that efficient reasoning depends not only on which visual information survives pruning, but also on how reliably the model learns to use it.

↑