arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30667cs.CVcs.LG

StarWM:基于自监督训练注意力路由的鲁棒世界模型

StarWM: Self-Supervised Trained Attention Routing for Robust World Models

Zeqiang Zhang, Fabian Wurzberger, Maximilian Otte, Daniel Schmid, Sebastian Gottwald, Arne Peter Raulf, Daniel Alexander Braun

首次发表
浏览论文内容

中文总结 AI 辅助

StarWM通过自监督训练的交叉注意力路由决定重建区域,结合双流解码器与停止梯度,在动态视频干扰下实现鲁棒世界模型,兼顾动态保真与干扰剔除。

中文摘要 AI 辅助

一个鲁棒的世界模型必须在忠实捕捉环境动态与抽象掉无关内容之间取得平衡。虽然基于重建的世界模型确保了忠实的监督,但它们在视觉任务中按像素面积而非动态相关性分配表征容量,这可能导致任务无关内容主导学习到的表征。相反,免重建方法避免了这种偏差,但存在丢弃可能相关信息的风�险。我们提出StarWM,它使用一个在自监督动态上训练的交叉注意力模块来决定重建应用于何处。一个双流解码器随后将重建限制在注意力区域,并通过停止梯度屏障防止两个目标之间的干扰。这些组件使得重建能够监督注意力区域的视觉内容,而不会用非预测信息污染潜在表征。在具有动态视频背景的DeepMind Control上,默认(无奖励)的StarWM在随机帧干扰下取得了最强性能,并在顺序视频下大幅优于基于重建的基线。此外,其奖励增强变体在顺序视频上达到或超过免重建方法,在所有干扰场景中取得了最高总体回报。机制探测证实,StarWM通过长视界想象以近乎完美的保真度保留状态属性,同时系统性地丢弃干扰物。

英文摘要

A robust world model must strike the balance between faithfully capturing environmental dynamics and abstracting away from irrelevant content. While reconstruction-based world models ensure faithful supervision, they misallocate representational capacity by pixel area rather than dynamics relevance for visual tasks, which can cause task-irrelevant content to dominate the learned representation. Alternatively, reconstruction-free methods avoid this bias but risk discarding possibly relevant information. We propose StarWM, which uses a cross-attention module trained on self-supervised dynamics to decide where reconstruction applies. A dual-stream decoder then restricts reconstruction to the attended regions, with stop-gradient barriers preventing interference between the two objectives. These components allows reconstruction to supervise the visual content of attended regions without contaminating the latent with non-predictive information. On DeepMind Control with dynamic video backgrounds, default (reward-free) StarWM achieves the strongest performance under random-frame distractors and substantially outperforms reconstruction-based baselines under sequential video. In addition, its reward-augmented variant matches or exceeds reconstruction-free methods on sequential video, achieving the highest overall return across all distractor regimes. Mechanistic probing confirms StarWM preserves state attributes with near-perfect fidelity through long-horizon imagination while systematically discarding distractors.

发表机构

  • German Aerospace Center (DLR)(德国航空航天中心(DLR))
  • Ulm University(乌尔姆大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑