arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33218cs.CVcs.AIcs.RO

Scope-WM:用于高效视觉世界模型的作用域计算

Scope-WM: Scoped Computation for Efficient Visual World Models

Chunzheng Li, Zesheng Jia, Hongda Zhang, Jiaying Tang, Yuntian Wang, Siao Liu, Jin Wang

AI总结:

Scope-WM通过作用域计算,将动态预测集中于相关令牌并重用优质动作序列,实现高效视觉世界模型,显著降低内存和规划时间,同时保持性能。

AI中文摘要:

视觉世界模型通过预测未来观测来实现机器人规划,但密集的潜在状态传播和样本密集的轨迹优化会导致高推理延迟和峰值内存使用,限制了在资源受限平台上的实时部署。现有的稀疏世界模型加速方法要么依赖无引导的令牌稀疏化,这可能丢弃与规划相关的信息并限制可实现的稀疏性,要么引入重型辅助模块和繁琐的多阶段训练流程。在这项工作中,我们提出了Scope-WM,一种高效的视觉世界模型,它将计算作用域限定到与预测相关的潜在区域和有前景的动作序列。Scope-WM将预测相关性提炼到一个轻量级的动作条件选择器中,并仅对选定的紧凑令牌子集执行完整动态预测。它使用前景状态及其变化的紧凑摘要更新剩余令牌,使背景能够感知前景动态,而无需昂贵的令牌间交互。在规划过程中,Scope-WM保留并重用初始MPC搜索中发现的高质量动作序列,在减少的展开预算下聚焦后续搜索。由此产生的流程仅需要一次性的选择器蒸馏,随后对稀疏世界模型进行一次联合训练阶段。在具有挑战性的Push-T任务上,Scope-WM将峰值GPU内存使用量和规划时间分别降低到密集DINO-WM的18.1%和14.3%,对应6.97倍的规划加速,同时保持具有竞争力的任务性能。在五个不同的视觉规划任务上的进一步评估证明了Scope-WM的普遍适用性。代码可在以下网址获取:此https URL。

英文摘要:

Visual world models enable robotic planning by predicting future observations, but dense latent-state propagation and sample-intensive trajectory optimization incur high inference latency and peak memory usage, limiting real-time deployment on resource-constrained platforms. Existing sparse world-model acceleration methods either rely on unguided token sparsification, which may discard planning-relevant information and restrict achievable sparsity, or introduce heavy auxiliary modules and cumbersome multi-stage training pipelines. In this work, we present Scope-WM, an efficient visual world model that scopes computation to prediction-relevant latent regions and promising action sequences. Scope-WM distills prediction relevance into a lightweight action-conditioned selector and applies full dynamics prediction only to a compact subset of selected tokens. It updates the remaining tokens using a compact summary of foreground states and their changes, allowing the background to perceive foreground dynamics without costly token-to-token interactions. During planning, Scope-WM preserves and reuses high-quality action sequences discovered during the initial MPC search, focusing subsequent search under reduced rollout budgets. The resulting pipeline requires only a one-off selector distillation followed by a single joint training stage for the sparse world model. On the challenging Push-T task, Scope-WM reduces peak GPU memory usage and planning time to $18.1\%$ and $14.3\%$ of those of dense DINO-WM, respectively, corresponding to a $6.97\times$ planning speedup, while maintaining competitive task performance. Further evaluations across five diverse visual planning tasks demonstrate the general applicability of Scope-WM. Code is available at https://github.com/ChunZheng2022/Scope-WM.

↑