AI 中文总结
提出LT-OPD框架,通过在线策略自蒸馏和预算课程,在极端视觉令牌缩减下恢复MLLM性能,平均保留性能从68.6%提升至82.3%,并显著降低KV缓存和预填充计算量。
AI 中文摘要
视觉令牌缩减是加速多模态大语言模型(MLLMs)的有效方式,但在极低令牌预算下性能会迅速恶化。现有工作已探索了视觉令牌选择以及基于训练的自适应缩减视觉输入的方法。我们更进一步,探究一个重度压缩的MLLM应如何从其自身生成所诱导的状态中学习。这一设定自然要求采用在线策略自蒸馏:一个重度压缩的模型在其自身生成所诱导的状态上接受监督,而其全令牌对应版本则充当信息丰富的教师。基于这一见解,我们提出了LT-OPD,一个面向极端视觉令牌缩减的训练框架。学生模型仅使用一小部分视觉令牌生成响应,而同一MLLM的冻结全令牌副本则沿这些学生生成的轨迹提供分布监督。为了在视觉证据严重受限时稳定在线策略学习,我们进一步引入了预算级课程,在训练过程中逐步降低令牌预算。在Qwen3.5-4B上的九个基准测试中,LT-OPD将5%视觉令牌保留下的平均保留性能从68.6%提升至82.3%,在相同预算下优于免训练、基于训练和强化学习的基线方法。这些增益一致地迁移到Qwen3.5-9B、GLM-4.6V-9B和LLaVA-OV-1.5-4B上。LT-OPD还将KV缓存使用量减少了85.2%,预填充FLOPs减少了85.4%,且无额外推理开销,表明在线策略学习能够大幅恢复因极端视觉令牌缩减而损失的能力。
英文摘要
Visual token reduction is an effective way to accelerate multimodal large language models (MLLMs), but performance deteriorates rapidly under extremely low token budgets. Existing work has explored both visual-token selection and training-based adaptation to reduced visual inputs. We take a step further by asking how a heavily compressed MLLM should learn from the states induced by its own generations. This setting naturally calls for on-policy self-distillation: a heavily compressed model is supervised on the states induced by its own generations, while its full-token counterpart serves as an information-rich teacher. Based on this insight, we propose LT-OPD, a training framework for extreme visual-token reduction. The student rolls out responses with only a small fraction of visual tokens, and a frozen full-token copy of the same MLLM provides distributional supervision along these student-generated trajectories. To stabilize on-policy learning when visual evidence is severely limited, we further introduce a budget-level curriculum that progressively decreases the token budget during training. Across nine benchmarks on Qwen3.5-4B, LT-OPD raises average retained performance under 5% visual-token retention from 68.6% to 82.3%, outperforming training-free, training-based, and reinforcement-learning baselines at the same budget. The gains transfer consistently to Qwen3.5-9B, GLM-4.6V-9B, and LLaVA-OV-1.5-4B. LT-OPD also reduces KV-cache usage by 85.2% and prefill FLOPs by 85.4% without additional inference overhead, demonstrating that on-policy learning can substantially recover capabilities lost to extreme visual-token reduction.
CommentsCode is at https://github.com/Yrxxxxxxxx1007/LT-OPD