arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更少数据,更好时机:时间视频定位中VLM在线策略蒸馏的学生-课程耦合

Less Data, Better Timing: Student-Curriculum Coupling for VLM On-Policy Distillation in Temporal Video Grounding

Jiacheng Qiu, Yunsoo Kim, Ruichen Xu, Jian Luo, Petar M. Djurić, Sima Mofakham

arXiv 2609.40055首次发表:更新:

发表机构

State University of New York at Stony Brook(纽约州立大学石溪分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对时间视频定位中在线策略蒸馏监督价值随时间变化的问题,提出学生-课程耦合框架,动态调整监督子集,在减少60%数据、50.4%训练时间下提升平均召回率5.1%。

AI 中文摘要

在线策略蒸馏(OPD)直接在学生生成的轨迹上提供密集监督,使其成为时间视频定位(TVG)中视觉语言模型有效的后训练策略。然而,现有流程通常从固定的教师模型和初始学生状态构建训练课程,隐含假设所选示例在整个优化过程中保持正面的监督价值。我们表明,监督可信度与监督必要性是不同但耦合的:前者关注目标的可信性,而后者随学生当前任务能力而变化;两者共同塑造监督价值。基于这一耦合视角,我们引入了学生-课程耦合(SCC),一个闭环框架,其中紧凑的锚点-前沿课程定义候选监督空间,而不断演进的学生动态确定其活跃子集。因此,监督可以根据能力变化被激活、暂停或重新激活,将教师计算和优化集中在当前任务层面的缺陷上。在三个TVG基准上,SCC在原始课程上相比Video-OPD实现了平均召回率5.1%的相对提升,同时使用60.0%更少的训练示例,并将训练时间减少50.4%。消融研究支持了能力结构化的课程设计和学生依赖的监督在实现这些增益中的互补作用。总之,这些结果确立了SCC作为TVG后训练的数据和计算高效框架,通过将可信监督与学生不断演进的学习需求对齐,实现更强的时间定位。

英文摘要

On-policy distillation (OPD) provides dense supervision directly on student-generated trajectories, making it an effective post-training strategy for vision-language models in temporal video grounding (TVG). However, existing pipelines typically construct the training curriculum from a fixed teacher and the initial student state, implicitly assuming that selected examples retain positive supervision value throughout optimization. We show that supervision trustworthiness and supervision necessity are distinct yet coupled: the former concerns target credibility, while the latter varies with the student's current task competence; together, they shape supervision value. Building on this coupled view, we introduce Student-Curriculum Coupling (SCC), a closed-loop framework in which a compact Anchor-Frontier curriculum defines the candidate supervision space and the evolving student dynamically determines its active subset. Supervision can therefore be activated, suspended, or reactivated as competence changes, concentrating teacher computation and optimization on current task-level deficits. Across three TVG benchmarks, SCC achieves a 5.1% relative improvement in mean recall over Video-OPD on its original curriculum, while using 60.0% fewer training examples and reducing training time by 50.4%. Ablations support the complementary roles of capability-structured curriculum design and student-dependent supervision in achieving these gains. Together, these results establish SCC as a data- and compute-efficient framework for TVG post-training, delivering stronger temporal grounding by aligning trustworthy supervision with the student's evolving learning needs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑