arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.25891cs.CVcs.AI

BAS-OPD:面向细粒度多模态感知的预算感知选择性在线策略自蒸馏

BAS-OPD: Budget-Aware Selective On-Policy Self-Distillation for Fine-Grained Multimodal Perception

  • HFIPS, Chinese Academy of Sciences(中国科学院合肥物质科学研究院)
  • University of Science and Technology of China(中国科学技术大学)
  • University of California, Los Angeles(加利福尼亚大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

Zihan Chen, Hengguang Zhou, Yuan Kang, Yiming Zhang, Wenhui Fang, Zenghui Ding, Yining Sun, Cho-Jui Hsieh

AI总结:

针对多模态大模型细粒度感知中教师查询成本高的问题,提出预算感知的选择性在线策略自蒸馏框架BAS-OPD,通过智能选择信息样本分配监督,在保持性能的同时大幅降低训练成本。

AI中文摘要:

多模态大语言模型(MLLMs)在处理完整图像时,常常难以进行细粒度的视觉感知,因为关键证据可能仅出现在局部区域。在线策略自蒸馏(OPD)能够将信息丰富的视图中的特权视觉知识迁移到全图像策略中,但为每次轨迹查询教师模型会带来大量的监督成本。在这项工作中,我们提出了BAS-OPD,一种预算感知的选择性OPD框架,在有限的查询预算下分配教师监督。BAS-OPD并非查询所有轨迹,而是选择信息丰富的样本,同时保持全批次的学生生成。我们探索了随机、基于不确定性和基于学习效用的选择策略,其中学习的选择器从分离的轨迹统计量和在线效用信号(由学生-教师一致性和教师置信度导出)中估计查询价值,无需额外的学生前向传播。BAS-OPD仅改变训练时的监督分配,并保持单次前向的全图像推理。在细粒度多模态感知基准上的实验表明,BAS-OPD在显著降低教师监督成本的同时取得了强劲的性能,突显了在受限预算下选择性OPD的有效性。

英文摘要:

Multimodal large language models (MLLMs) often struggle with fine-grained visual perception when processing complete images, as critical evidence may only appear in local regions. On-policy self-distillation (OPD) enables transferring privileged visual knowledge from informative views to full-image policies, but querying the teacher for every rollout introduces substantial supervision costs. In this work, we propose BAS-OPD, a budget-aware selective OPD framework that allocates teacher supervision under limited query budgets. Instead of querying all rollouts, BAS-OPD selects informative samples while maintaining full-batch student generation. We explore random, uncertainty-based, and learned utility-based selection strategies, where the learned selector estimates query value from detached rollout statistics and online utility signals derived from student--teacher agreement and teacher confidence without additional student forward passes. BAS-OPD only changes training-time supervision allocation and preserves single-pass full-image inference. Experiments on fine-grained multimodal perception benchmarks demonstrate that BAS-OPD achieves strong performance while substantially reducing teacher supervision costs, highlighting the effectiveness of selective OPD under constrained budgets.

↑