TaskSense:聚焦世界模型中的关键内容
TaskSense: Focusing on What Matters in World Models
浏览论文内容
中文总结 AI 辅助
TaskSense是一种以任务为中心的世界建模框架,通过可微随机空间注意力机制结合辅助逆动力学目标,提升了视觉控制模型对干扰环境的鲁棒性,在Distracting Control Suite上性能优于DreamerV3。
中文摘要 AI 辅助
用于视觉控制的世界模型通常通过重构观测结果来学习紧凑的隐状态,这会隐性地鼓励表示保留整个视觉输入的信息。然而,与任务相关的内容往往仅占观测的一小部分,而背景杂波和干扰项会消耗宝贵的表示容量。视觉重构与控制目标之间的这种不匹配,会使隐表示偏向于对与任务无关的视觉内容进行建模,从而削弱与控制相关特征的学习信号,并在视觉干扰下严重降低下游性能。我们提出了TaskSense,一种以任务为中心的世界建模框架,该框架通过基于前一个隐状态的可微随机空间注意力机制,在隐编码前强制执行任务相关性。为了引导注意力转向与控制相关的区域,我们用辅助逆动力学目标来增强训练。该世界模型不再重构完整观测,而是仅重构被关注的区域,这鼓励隐表示保留与任务相关的信息,同时丢弃无关的视觉内容。解码器还以采样的注意力图为条件,尽管注意力是随机的,仍能实现一致的重构。与DreamerV3基线相比,TaskSense在DeepMind Control Suite上保持了有竞争力的性能,同时在Distracting Control Suite上始终优于DreamerV3,表现出对视觉干扰显著提升的鲁棒性。定性分析进一步证实,在逆动力学监督的引导下,学习到的注意力会持续定位与控制相关的区域,同时抑制无关的视觉内容。
英文摘要
World models for visual control typically learn compact latent states by reconstructing observations, implicitly encouraging representations to preserve information across the entire visual input. However, task-relevant content often occupies only a small fraction of the observation, while background clutter and distractors consume valuable representational capacity. This mismatch between visual reconstruction and control objectives biases latent representations to model task-irrelevant visual content, diluting learning signals for control-relevant features and severely degrading downstream performance under visual distractions. We introduce TaskSense, a task-centric world modeling framework that enforces task relevance before latent encoding through a differentiable stochastic spatial attention mechanism conditioned on the previous latent state. To steer attention toward control-relevant regions, we augment training with an auxiliary inverse-dynamics objective. Rather than reconstructing the full observation, the world model reconstructs only the attended regions, encouraging latent representations to preserve task-relevant information while discarding irrelevant visual content. The decoder is further conditioned on the sampled attention map, enabling consistent reconstruction despite stochastic attention. Compared with the DreamerV3 baseline, TaskSense maintains competitive performance on the DeepMind Control Suite while consistently outperforming DreamerV3 on the Distracting Control Suite, demonstrating substantially improved robustness to visual distractions. Qualitative analysis further confirms that the learned attention, guided by inverse-dynamics supervision, consistently localizes control-relevant regions while suppressing irrelevant visual content.