arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.26902cs.CV

绑定主体,释放场景:面向长时序自回归视频生成的查询感知记忆路由

Tether the Subject, Release the Scene: Query-Aware Memory Routing for Long-Horizon Autoregressive Video Generation

Chen Li, Peng Zhang, Hanyu Zhou, Jialong Zuo, Fei Wang, Daiguo Zhou, Nong Sang, Changxin Gao

首次发表
浏览论文内容

中文总结 AI 辅助

针对流式自回归视频生成中记忆锚定的场景欠进展问题,提出无需训练的查询感知时空记忆路由器 TetherMem,分离主体与场景查询调节历史访问,在长视频生成的整体质量与场景进展指标上优于基线方法。

中文摘要 AI 辅助

流式自回归视频模型通过逐块生成长视频,利用历史记忆维持一致性。现有方法通常以相似策略将主体与场景查询暴露给历史,这虽能稳定主体,但也会将背景、视角和场景结构锁定在之前生成的状态,即便局部运动仍在继续。我们将此失效称为“记忆锚定的场景欠进展”,仅靠一致性和运动指标可能无法检测到它。我们提出 TetherMem,一种用于冻结视频生成器的无需训练的查询感知时空记忆路由器。TetherMem 分离主体与场景查询,并通过基于区域和年龄的先验调节历史访问:主体查询保留承载身份的历史,而场景查询减少对主体历史和陈旧背景的依赖。在来自 10 名标注者的 2400 次盲选成对判断中,TetherMem 在 8 个流式长视频基线中,整体质量(0.780)和场景进展(0.769)的估计期望偏好均最高。在完整的 30 秒视频上,它能在保留主体可识别性和时间连续性的同时,维持背景、视角和场景状态的变化。

英文摘要

Streaming autoregressive video models generate long videos chunk by chunk, using historical memory to maintain consistency. Existing methods typically expose subject and scene queries to history through similar policies. This stabilizes the subject, but can also lock backgrounds, viewpoints, and scene structure to previously generated states even when local motion continues. We call this failure memory-anchored scene under-progression; consistency and motion metrics alone can miss it. We introduce TetherMem, a training-free, query-aware spatiotemporal memory router for frozen video generators. TetherMem separates subject and scene queries and modulates historical access with region- and age-conditioned priors: subject queries retain identity-bearing history, while scene queries reduce reliance on subject history and stale backgrounds. Across 2,400 blinded pairwise judgments from 10 annotators, TetherMem achieves the highest estimated expected preference among eight streaming long-video baselines for overall quality (0.780) and scene progression (0.769). On complete 30-second videos, it sustains changes in background, viewpoint, and scene state while preserving subject recognizability and temporal continuity.

发表机构

  • Huazhong University of Science and Technology(华中科技大学)
  • MiLM Plus, Xiaomi Inc.(小米公司MiLM Plus)

机构由 AI 辅助整理,请以论文原文为准。

↑