发表机构
Mohamed bin Zayed University of Artificial Intelligence; ELLIS Institute Finland; Aalto University(穆罕默德·本·扎耶德人工智能大学; 芬兰ELLIS研究所; 阿尔托大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出首个完全自包含的连续视频状态跟踪框架S$^3$T,利用时间采样密度作为特权信息实现无监督训练,在LLaVA-OneVision-2-8B等基准上显著提升VSTAT等任务性能,且可迁移至真实视频。
AI 中文摘要
我们提出S$^3$T(Self-Supervised Self-Distillation over Time,时间自监督自蒸馏),据我们所知,这是首个完全自包含的连续视频状态跟踪框架。该方法将时间采样密度视为特权信息,基于同一视频片段更密集的视角能更准确地恢复运行状态的假设,此视角作为教师模型,而权重相同的稀疏视角学生模型学习匹配其下一个标记分布。模型生成自身目标,因此训练无需标签、独立教师或奖励信号,且不增加推理成本。在LLaVA-OneVision-2-8B上,S$^3$T作为单一模型将VSTAT准确率提升1.74,通过souping提升2.38,通过额外视觉编码器适配提升2.70,而现有自演化方法几乎未改变状态跟踪性能。从未标记合成片段中学到的能力可迁移至真实视频,在VSTAT-YouTube状态跟踪问题上提升7.95,在MVBench动作计数上提升4.50。
英文摘要
We introduce S$^3$T (Self-Supervised Self-Distillation over Time), which, to the best of our knowledge, is the first fully self-contained framework for continuous video state tracking. Our method treats temporal sampling density as privileged information, based on the hypothesis that a denser view of the same clip recovers the running state more accurately. This view serves as the teacher, while a sparse-view student with the same weights learns to match its next-token distribution. The model generates its own target, so training requires no labels, separate teacher, or reward signal, and adds no inference cost. On LLaVA-OneVision-2-8B, S$^3$T improves VSTAT accuracy by $+1.74$ as a single model, $+2.38$ with souping, and $+2.70$ with additional vision-encoder adaptation, while prior self-evolving methods leave state tracking largely unchanged. The capability learned from unlabeled synthetic clips transfers to real videos, improving performance by $+7.95$ on VSTAT-YouTube state-tracking questions and $+4.50$ on MVBench Action Count.