arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14144cs.CVcs.AI

自监督视觉在线策略蒸馏

Self-Supervised Visual On-Policy Distillation

  • University of Maryland, College Park(马里兰大学帕克分校)
  • Georgia Institute of Technology(佐治亚理工学院)
  • Johns Hopkins University(约翰斯·霍普金斯大学)
  • University of Oxford(牛津大学)
  • UC San Diego(加利福尼亚大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

Yijiang Li, Yijun Liang, Yunjie Tian, Bingyang Wang, Ke Zhang, Zhenfei Yin, Di Fu, Philip Torr, Nuno Vasconcelos

AI总结:

该研究提出S²VOPD方法,通过从学生模型减信息而非给教师加特权信息,实现无特权信息的视觉在线策略蒸馏,在6个细粒度感知基准上显著提升Qwen3.5-4B性能,表现优于多数模型。

AI中文摘要:

视觉在线策略蒸馏高度依赖于信息丰富的教师-学生不对称性,这种不对称性要么来自更大、更强的教师模型,要么来自特权监督,如参考答案或感兴趣区域的真值标注。这提出了一个基本问题:当没有任何特权信息可用时,信息丰富的不对称性从何而来?我们通过反转不对称性的来源来回答这个问题。我们没有向教师添加特权信息,而是从学生中减去信息。这种不对称性免费创造了与教师拥有学生无法访问的信息时相同的有效学习信号,无需真值标注、奖励或单独的更强教师模型。基于这一原理,我们引入了自监督视觉在线策略蒸馏(S²VOPD),这是一种简单而有效的方法,可从不对称增强视图构建在线策略学习信号。S²VOPD将原始图像条件下的教师在线策略分布蒸馏为同一图像的强增强视图条件下的学生分布。我们系统地探索了广泛的视觉增强设计空间,发现:(1)不对称性很重要:所有四个增强族都能提高性能,而对称自蒸馏会降低性能;(2)强度很重要:性能在中等强度时达到峰值;(3)差距必须保持任务一致性:完全去除与问题相关证据的增强会导致大但无信息的差异。在六个细粒度感知基准上,S²VOPD将Qwen3.5-4B的性能从70.7%提高到77.4%,优于所有比较的开源模型,达到了2350亿参数的Qwen3-VL的水平,并且超过了GPT-5.4。在保持训练数据相同的情况下,它恢复了使用特权信息的方法所实现的96%的改进。

英文摘要:

Visual on-policy distillation relies heavily on an informative teacher-student asymmetry, through either a larger, stronger teacher or privileged supervision, such as reference answers or ground-truth regions of interest. This raises a fundamental question: where can informative asymmetry come from when nothing privileged is available? We answer this by inverting where the asymmetry comes from. Rather than adding privileged information to the teacher, we subtract information from the student. This asymmetry creates the same effective learning signal for free as a teacher with access to information unavailable to the student, without ground-truth annotations, rewards, or a separate stronger teacher model. Building on this principle, we introduce Self-Supervised Visual On-Policy Distillation (S$^2$VOPD), a simple yet effective method that constructs on-policy learning signals from asymmetric augmented views. S$^2$VOPD distills the teacher's distribution conditioned on the original image on-policy into the student distribution conditioned on a strongly augmented view of the same image. We systematically explore a broad design space of visual augmentations and uncover that (1) asymmetry matters: all four augmentation families improve performance, while symmetric self-distillation degrades it; (2) strength matters: performance peaks at a moderate strength; and (3) the gap must remain task-consistent: augmentations that completely remove the question-relevant evidence can induce large but uninformative discrepancies. Across six fine-grained perception benchmarks, S$^2$VOPD improves Qwen3.5-4B from 70.7% to 77.4%, above all open-source models compared, up to Qwen3-VL at 235B, and surpasses GPT-5.4. While holding training data the same, it recovers 96% of the improvement achieved by methods with privileged information. Website is at https://williamium3000.github.io/s2vopd

↑