arXivDaily arXiv每日学术速递 周一至周五更新
arXiv 2607.21848cs.CV

闭环:自回归生成渲染的无训练重访一致性

Closing the Loop: Training-Free Revisit Consistency for Autoregressive Generative Rendering

  • Roblox(罗布乐思)
  • The Pennsylvania State University(宾夕法尼亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

Wenchao Ma, Changran Liu, Sharon X. Huang, Haomiao Jiang

AI总结:

研究自回归生成渲染中相机重访位置时的不一致问题,利用3D引擎提供的时间和空间对应关系,将姿态匹配的历史潜在块检索到KV缓存,并使注意力偏向对应区域,在相关数据集上实验,提升了重访一致性且不损失视频质量。

AI中文摘要:

近期的条件视频生成模型在将3D引擎渲染(如深度图和无纹理几何体)转换为用于游戏和沉浸式内容创作的逼真视频方面展现出潜力。这些应用需要长期的自回归生成,在保留持久3D世界的同时持续合成新帧。自回归生成器使用有限的KV缓存逐块合成视频,当相机重访已逐出上下文的位置时,模型常生成不一致外观。本文通过利用3D引擎提供的对应关系解决此重访不一致问题,无需任何训练后处理。通过时间对应将姿态匹配的历史潜在块检索到KV缓存作为闭环记忆,空间对应通过相机姿态和深度重投影使令牌级注意力偏向检索块的几何对应区域。在从TartanAir和TartanGround数据集中挖掘的闭环轨迹上进行实验,结果表明该方法在不损失整体视频质量的情况下,重访一致性优于现有的无训练基线。

英文摘要:

Recent conditional video generation models have shown promising potentials to transform 3D engine renderings, such as depth maps and untextured geometry, into photorealistic videos for gaming and immersive content creation. These applications require long-horizon auto-regressive generation that continuously synthesizes new frames while preserving a persistent 3D world. Auto-regressive generators synthesize video chunk by chunk with a bounded KV cache, so when the camera revisits a location after its context has been evicted, the model often regenerates inconsistent appearance, even though the conditioning renderings (e.g., depth) remain perfectly aligned with the underlying geometry. We address this revisit inconsistency without any post-training by exploiting correspondences the 3D engine already provides: temporal correspondence retrieves pose-matched historical latent chunks into the KV cache as loop-closure memory, while spatial correspondence from camera pose and depth reprojection biases token-level attention toward geometrically corresponding regions of the retrieved chunks. We demonstrate our method on loop-closure trajectories mined from TartanAir and TartanGround dataset to mirror complicate real-world application scenarios, where it outperforms existing training-free baselines on revisit consistency without losing overall video quality. Project Page: https://wenchao-m.github.io/ClosetheLoop.github.io/

↑