arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.26657cs.ROcs.CV

Enfold:将世界生成器计算融入预测表示以实现高效具身控制

Enfold: Folding World Model Imagination into Predictive Representations for Ultra-Efficient Embodied Control

Weili Zeng, Yitong Xing, Fulong Liu, Chengqun Yang, Antao Xiang, Feng Tian, Jingnan Gao, Jisong Cai, Xin Wang, Xiaomin Wu, Yao Mu, Xiaokang Yang, Yichao Yan

首次发表
浏览论文内容

中文总结 AI 辅助

Enfold将世界生成器计算融入预测表示,在LIBERO等多任务中实现高效具身控制,大幅降低动作延迟,且能适应场景变化,为具身控制提供了新范式。

中文摘要 AI 辅助

世界生成模型通常通过其输出发挥作用:渲染的未来、基于视频的动作,或由成本高昂的生成分支计算出的潜在上下文。我们认为,其更具可复用性的资产是构建未来的计算过程。当生成器将受损的未来转化为连贯轨迹时,其中间状态会在不同抽象层级上组织外观、空间布局与交互。能否将这种未来生成计算内化于仅从当前推断出的表示中?我们提出Enfold,它将该计算转移至从当前视觉上下文和语言指令预测出的表示中。训练期间,生成器处理观测到的未来时暴露的多层级状态会监督仅使用当前数据的编码器。学习到的表示被反馈以条件未来生成,并被任务头读取,同时不允许任务梯度重塑编码器。部署时,动作预测不再执行生成器。在LIBERO、RoboTwin2.0及真实机器人任务中,Enfold在实现强控制的同时,相较于Fast-WAM将动作延迟降低了3.7倍,Enfold-FFlash则达到了10.1倍。表示分析显示,它抑制了无关变化,优先捕捉更长时间范围内出现的变化。当当前场景被人为干预改变时,生成的后续内容与执行的动作均会适应,这与固定轨迹回放不符。这些结果将世界生成器重新定义为预测控制表示的来源:若其内部结构可融入当前,未来无需在每一步都被具象化。

英文摘要

World generative models are typically used through what they produce: a rendered future, a video-conditioned action, or latent context computed by a costly generative branch. We argue that their more reusable asset is the computation that constructs a future. As a generator transforms a corrupted future into a coherent trajectory, its intermediate states organize appearance, spatial layout, and interaction across levels of abstraction. Can this future-generative computation be internalized in a representation inferred from the present alone? We present Enfold, which transfers this computation into a representation predicted from the current visual context and language instruction. During training, multi-level states exposed as the generator processes the observed future supervise a current-only encoder. The learned representation is fed back to condition future generation and is read by task heads without allowing task gradients to reshape the encoder. At deployment, action prediction no longer executes the generator. Across LIBERO, RoboTwin2.0, and real-robot tasks, Enfold supports strong control while reducing action latency by $3.7\times$ relative to Fast--WAM, Enfold-Flash reaches $10.1\times$. Representation analyses show that it suppresses nuisance variation and preferentially captures changes that emerge over longer horizons. When the current scene is altered by human intervention, both the generated continuation and the executed actions adapt, which is inconsistent with fixed trajectory replay. These results recast a world generator as a source of predictive control representations: its future need not be materialized at every step if its internal structure can be enfolded into the present.

发表机构

  • Shanghai Jiao Tong University(上海交通大学)
  • South China University of Technology(华南理工大学)
  • Qilu University of Technology (Shandong Academy of Sciences)(齐鲁工业大学(山东省科学院))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑