arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

IntactWorld:基于完整特征的联合世界建模

IntactWorld: Joint World Modeling with Intact Features

Boming Tan, Xiangdong Zhang, Yan Xia, Qi Zhu, Deyi Ji, Xue Yang, Shaofeng Zhang

arXiv 2610.11174首次发表:更新:

发表机构

University of Science and Technology of China; Shanghai Jiao Tong University; KOKONI 3D, Moxin Technology(中国科学技术大学; 上海交通大学; KOKONI 3D、墨芯科技)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对现有视频生成模型世界建模时特征压缩导致结构信息损失的问题,提出IntactWorld架构,采用中间层预测干净特征、全量到紧凑训练范式,在VBench 2.0上优于基线2.46个点。

AI 中文摘要

尽管近期的视频生成模型能合成高度逼真的视觉内容,但它们缺乏对真实世界内在逻辑的真正理解。现有方法试图通过内化多样的世界知识来理解世界,但受计算开销或维度对齐的限制,其学习过程不可避免地会压缩特征,导致结构信息的严重损失。为解决这一问题,我们提出了IntactWorld,这是一种利用未压缩完整特征的联合世界建模架构。由于数据自然存在于高维空间中的低维流形上,预测该未压缩高维空间中的流速度v会引发严重的流形间隙。为成功消除这一优化瓶颈,我们的框架转而在中间层预测干净特征x₀。此外,为缓解纳入完整世界知识的计算开销,我们引入了全量到紧凑训练范式。该范式通过用高度精炼的CLS令牌替换原始全量特征,实现了高效的单分支引导,将空间内存消耗降低11.4%,并将推理延迟缩短43.8%。大量评估证明了IntactWorld的有效性,其在VBench 2.0基准上比现有基线高出2.46个点。

英文摘要

While recent video generation models synthesize highly realistic visuals, they lack a genuine understanding of intrinsic real-world logic. Existing methods attempt to understand the world by internalizing diverse world knowledge, yet constrained by computational overhead or dimensionality alignment, their learning processes inevitably compress features, causing a severe loss of structural information. To address this, we propose \textbf{IntactWorld}, a \textbf{Joint World Modeling Architecture} utilizing uncompressed \textbf{Intact Features}. Since data naturally reside on a low-dimensional manifold within a high-dimensional space, predicting the flow velocity $v$ within this uncompressed high-dimensional space induces a severe manifold gap. To successfully eliminate this optimization bottleneck, our framework instead predicts the clean feature $x_0$ at intermediate layers. Furthermore, to mitigate the computational overhead of incorporating complete world knowledge, we introduce a \textit{Full-to-Compact Training Paradigm}. By replacing raw full features with highly refined CLS tokens, this paradigm enables efficient single-branch guidance, reducing spatial memory consumption by 11.4\% and cutting inference latency by 43.8\%. Extensive evaluations demonstrate the effectiveness of IntactWorld, outperforming established baselines by 2.46 points on the VBench 2.0 benchmark.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑