发表机构
Sun Yat-sen University; The University of British Columbia; Vast Intelligence Lab(中山大学; 不列颠哥伦比亚大学; 威斯特智能实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出令牌预算蒸馏(TBD)框架,通过双路径师生设计结合LoRA适配器与FlashVID压缩,在固定令牌预算下适配视频VLMs,大幅压缩令牌时仍能保留高语义性能,在多基准测试中优于仅压缩基线。
AI 中文摘要
适配视频视觉语言模型(VLMs)的计算成本高昂,因为视频输入会产生大量视觉令牌,导致微调与推理均代价不菲。尽管视觉令牌压缩可降低该开销,但直接在压缩输入上适配常引发语义漂移与显著性能下降。本文提出令牌预算蒸馏(TBD),这是一种在固定令牌预算下适配视频VLMs的参数高效微调框架。TBD冻结预训练骨干网络,仅更新LoRA适配器,并将基于FlashVID的视觉令牌压缩整合至视频通路中。为在压缩下保留全令牌语义,TBD采用双路径师生设计:全令牌教师提供稳定监督,压缩学生则通过任务损失、答案区域KL蒸馏、真实值锚定间隔蒸馏及感知可靠性的知识蒸馏(KD)控制进行优化。该设计使学生模型在大幅减少令牌的同时,可恢复全令牌模型的语义行为。我们在LLaVA-Video、LLaVA-OneVision、Qwen3-VL-8B-Instruct三种视频VLM骨干网络及四个视频理解基准上评估TBD。在中等与大幅压缩场景下,TBD均持续优于仅压缩基线。在保留率R=10%时,LLaVA-Video上TBD保留了原始模型97.0%的平均准确率;在LLaVA-OneVision上R=10%时,其平均得分为58.4,达到100.0%的相对准确率。
英文摘要
Adapting video vision-language models (VLMs) is computationally expensive because video inputs produce a large number of visual tokens, making both fine-tuning and inference costly. Although visual token compression can reduce this overhead, direct adaptation on compressed inputs often causes semantic drift and noticeable performance degradation. We present Token-Budget Distillation (TBD), a parameter-efficient fine-tuning framework for adapting video VLMs under a fixed token budget. TBD freezes the pretrained backbone, updates only LoRA adapters, and integrates FlashVID-based visual token compression into the video pathway. To preserve full-token semantics under compression, TBD employs a dual-path teacher-student design, where a full-token teacher provides stable supervision and a compressed student is optimized with task loss, answer-region KL distillation, GT-anchored margin distillation, and reliability-aware KD control. This design enables the student to recover the semantic behavior of the full-token model while remaining efficient under aggressive token reduction. We evaluate TBD on three video VLM backbones, including LLaVA-Video, LLaVA-OneVision, and Qwen3-VL-8B-Instruct, across four video understanding benchmarks. TBD consistently outperforms compression-only baselines under both moderate and aggressive compression. On LLaVA-Video at retention ratio R = 10 percent, TBD preserves 97.0 percent of the Vanilla model's average accuracy; on LLaVA-OneVision at R = 10 percent, it achieves an average score of 58.4 and matches 100.0 percent relative accuracy.
CommentsACM MM 2026