发表机构
Baidu; University of Chinese Academy of Sciences; China University of Petroleum, Beijing(百度; 中国科学院大学; 中国石油大学(北京))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对流式视频理解中直接应用同策略自蒸馏导致的偏好错位和记忆崩溃问题,提出事件接地自蒸馏(EGSD),通过事件作为特权信息重加权教师并引入覆盖奖励,在主流基准上取得领先性能。
AI 中文摘要
实时视频理解需要增量式地维护流式内容的记忆,而优化这一过程需要密集的过程信号。同策略自蒸馏(OPSD)允许一个模型同时充当教师和学生,其中教师接收额外的特权信息(如问题和真实答案(GT)),能够提供此类令牌级信号。然而,将其直接应用于流式视频会引发两个问题。(1)学生无法进行端到端优化,因为记忆在问题到达之前就已写入,而教师却利用问题和真实答案的特权对其进行评分,导致两者的偏好不一致。(2)有效实体记忆崩溃,即问题和真实答案的特权使教师只偏好与问题相关的实体,而记忆上的令牌均值平均使得信号对记忆覆盖的实体数量不敏感,这两者都推动记忆偏离流式处理所需的多样性。为解决这些问题,我们提出了事件接地自蒸馏(EGSD),它将流式记忆表征为对可验证事件(关键视觉实体、动作和细节)的增量更新,并在此基础上针对上述两个问题。对于问题(1),我们将OPSD信号调整为与结果奖励相结合的乘法权重;对于问题(2),我们使用事件作为特权信息对教师进行重新加权,以抵消其问题相关性偏差,并添加实体覆盖奖励以提供令牌均值教师所缺乏的覆盖偏好。在主流在线和离线基准上的大量实验表明,EGSD取得了强劲的性能,在StreamingBench上达到79.8%,在OVO-Bench实时赛道上达到73.4%,而记忆分析显示,有效实体召回率在记忆长度仅增加6.8%的情况下提升了17.4%。
英文摘要
Real-time video understanding requires incrementally maintaining a memory of streaming content, and optimizing this requires dense process signals. On-Policy Self-Distillation (OPSD), which lets one model serve as both teacher and student with the teacher receiving additional privileged information such as the question and ground-truth (GT) answer, can supply such token-level signals. However, applying it directly to streaming video raises two problems. (1) The student cannot be optimized end-to-end, where memory is written before the question arrives, yet the teacher scores it with the question-and-GT privilege, misaligning their preferences. (2) Effective-entity memory collapses, where the question-and-GT privilege makes the teacher favor only question-relevant entities, and token-mean averaging over a memory renders its signal invariant to how many entities that memory covers, both driving memory against the streaming need for diversity. To address these issues, we propose Event-Grounded Self-Distillation (EGSD), which characterizes streaming memory as an incremental update over verifiable Events (key visual entities, actions, and details) and targets the two problems on this basis. For problem (1), we adapt the OPSD signal into a multiplicative weight combined with the outcome reward; for problem (2), we re-weight the teacher with Events as privileged information to counter its question-relevance bias, and add an entity-coverage reward to supply the coverage preference the token-mean teacher lacks. Extensive experiments on mainstream online and offline benchmarks show EGSD achieves strong performance, reaching 79.8% on StreamingBench and 73.4% on the OVO-Bench Real-Time track, while memory analysis shows effective-entity recall rises 17.4% at only 6.8% more memory length.