arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.37690cs.CV

Honeycomb:视频世界模型的恒定大小场景记忆表示

Honeycomb: Constant-Size Scene Memory Representation for Video World Models

Jack Wei Lun Shi, Kaichen Zhou, Haoyu Chen, Yufeng Weng, Keane Ong, Ruojin Cai, Hang Hua, Justin K. W. Yeoh, Mengyu Wang

首次发表
浏览论文内容

中文总结 AI 辅助

Honeycomb通过HexMemory固定大小记忆表示,实现视频世界模型的长时程生成一致性,避免存储增长,并在WorldScore和RealEstate10K上验证了高质量与稳健性。

中文摘要 AI 辅助

视频世界模型需要持久场景记忆以在长时程视频生成过程中保持一致性。现有的空间记忆系统累积RGB观测或潜在特征,导致存储需求随生成过程增长。我们引入Honeycomb,一种基于HexMemory的视频世界模型,HexMemory是一种紧凑的低秩表示,将场景特征存储在由六个空间和时空平面组成的固定大小记忆中。前馈写入器将每个新生成的视频块映射到平面特征。随着空间覆盖或时间范围扩大,HexMemory在保持维度不变的情况下扭曲现有平面,然后通过置信度加权池化和学习残差校正整合新特征。读取器从HexMemory检索潜在特征以条件化后续视频生成。由于写入器仅处理最新块中的观测,Honeycomb避免了逐场景优化和对完整生成历史的重复处理。在WorldScore和RealEstate10K上的实验证明了强视频生成质量和在重新访问先前观测区域时的稳健一致性,同时在整个生成过程中保持恒定的特征存储需求。代码和额外可视化可在我们的此https URL上获取。

英文摘要

Video world models require persistent scene memory to maintain consistency during long-horizon video generation. Existing spatial memories accumulate RGB observations or latent features, increasing storage requirements as generation proceeds. We introduce Honeycomb, a video world model built on HexMemory, our proposed low-rank representation for storing scene features in a fixed-size memory with a total of six spatial and spatiotemporal planes. A feed-forward writer maps each generated chunk into new plane features. As the spatial coverage or temporal range expands, we warp the previous planes while preserving their dimensions, then fuse them with the new features through confidence-weighted pooling and a learned residual correction. A reader retrieves latents from HexMemory to condition subsequent video generation. The writer processes only observations from the new chunk, avoiding per-scene optimization and repeated processing of the full history. Experiments on WorldScore and RealEstate10K demonstrate strong video generation quality and robust revisit consistency while keeping HexMemory feature storage constant throughout generation. Code and additional visualizations are available on our project page at https://jackswl.github.io/honeycomb/.

发表机构

  • Harvard University(哈佛大学)
  • National University of Singapore(新加坡国立大学)
  • MIT(麻省理工学院)
  • MIT-IBM Watson AI Lab(MIT-IBM沃森人工智能实验室)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

相关深度报道

↑