CA-OPD:面向结构化视觉预测的置信度感知在线策略蒸馏
CA-OPD: Confidence-Aware On-Policy Distillation for Structured Visual Prediction
浏览论文内容
中文总结 AI 辅助
针对自回归视觉语言模型的复合误差问题,提出CA-OPD框架,通过教师置信度实现自适应监督,在GUI定位等六基准上大幅提升Qwen3.5-0.8B性能。
中文摘要 AI 辅助
自回归视觉语言模型统一了异构感知任务,但极易出现复合误差。在线策略蒸馏(OPD)通过在学生模型自身的 rollout 结果上训练,解决了训练-推理不匹配问题。然而,学生模型的预测不可靠,尤其在训练初期,会破坏轨迹并降低教师监督的质量。尽管近期的交错蒸馏方法允许教师验证并替换学生 token,但它们主要依赖刚性排名指标而非教师的精确置信度,且忽略了干预决策如何为 token 级监督提供信息。为解决这一问题,我们提出置信度感知在线策略蒸馏(CA-OPD)框架,将可靠的 rollout 构建与自适应监督相结合。CA-OPD 利用教师置信度选择性修正不可靠的学生转移,通过严格到宽松的调度逐步将 rollout 控制权转移给学生。关键在于,CA-OPD 将知识转移与这些干预决策对齐:被修正的位置接收来自教师预测的直接交叉熵监督,而被保留的位置则受益于教师的完整预测分布。在 GUI 定位和光学字符识别的多教师设置中进行评估,CA-OPD 在全部六个目标基准上大幅提升了 Qwen3.5-0.8B 基线,包括在 ScreenSpot-Pro 上提升 9.50 个点,在 OCRBench-v2 English 上提升 6.72 个点。受控研究进一步表明,提升取决于干预位置、渐进式 rollout 控制以及与干预对齐的监督,而非仅干预频率。
英文摘要
Autoregressive vision language models unify heterogeneous perception tasks but are highly susceptible to compounding errors. On-policy distillation (OPD) bridges the training-inference mismatch by training students on their own rollouts. However, unreliable student predictions, especially early in training, can derail the trajectory and degrade the quality of teacher supervision. While recent interleaved distillation methods allow the teacher to verify and replace student tokens, they primarily rely on rigid ranking metrics rather than exact teacher confidence, and they overlook how intervention decisions can inform token-level supervision. To address this, we introduce Confidence-Aware On-Policy Distillation (CA-OPD), a framework that couples reliable rollout construction with adaptive supervision. CA-OPD utilizes teacher confidence to selectively correct unreliable student transitions, gradually transferring rollout control to the student via a strict-to-relaxed schedule. Crucially, CA-OPD aligns knowledge transfer with these intervention decisions: corrected positions receive direct cross-entropy supervision from the teacher's prediction, while retained positions benefit from the teacher's full predictive distribution. Evaluated in a multi-teacher setting for GUI grounding and optical character recognition, CA-OPD substantially improves the Qwen3.5-0.8B baseline across all six target benchmarks, including gains of $9.50$ points on ScreenSpot-Pro and $6.72$ points on OCRBench-v2 English. Controlled studies further show that the gains depend on intervention placement, progressive rollout control, and intervention-aligned supervision, rather than intervention frequency alone.
发表机构
- Tianjin University(天津大学)
- Shanghai Jiao Tong University(上海交通大学)
- StepX(思必得科技(StepX))
- University of Science and Technology of China(中国科学技术大学)
机构由 AI 辅助整理,请以论文原文为准。