AI 中文总结
AtlasVLA是一种新型VLA模型框架,通过双记忆架构实现持久化世界-自我状态建模,在仅用腕部单目相机的情况下,在LIBERO等基准及真实世界长时序任务中性能超越多视图基线。
AI 中文摘要
尽管视觉-语言-动作(VLA)模型在具身智能领域取得了进展,但其本质上的反应式范式严重限制了在部分可观测和长时序任务中的性能。当仅使用腕部安装的单目相机时,这些模型不可避免地会因物体离开视野而产生感知遗忘,以及在多步骤执行过程中出现时间任务进度遗忘。为克服这些瓶颈,我们提出了AtlasVLA,这是一种新型框架,通过持久化世界-自我状态,从直接反应式操作转向主动推理。AtlasVLA具有双记忆架构:4D持久化世界状态记忆,将瞬时2D观测提升为全局更新的体素哈希空间状态,以解决视觉盲区;以及自我工作状态记忆,用于跟踪历史自我状态和任务进度。通过将扩散Transformer(DiT)以联合世界-自我状态为条件,AtlasVLA实现了稳健的空间推理。在LIBERO、RLBench和真实世界基准上的广泛评估表明,AtlasVLA仅使用单目腕部相机就达到了最先进的性能。值得注意的是,它显著优于多视图基线,在LIBERO-Long上实现了9.4%的绝对成功率提升,在真实世界长时序任务中实现了17.5%的绝对成功率提升。
英文摘要
While Vision-Language-Action (VLA) models have advanced embodied AI, their fundamentally reactive paradigm severely limits performance in partially observable and long-horizon tasks. When restricted to a single wrist-mounted camera, they inevitably suffer from perception forgetting as objects exit the field of view, and temporal task-progress forgetting} during multi-step execution. To overcome these bottlenecks, we propose AtlasVLA, a novel framework that transitions from direct reactive manipulation to proactive reasoning through a persistent world-ego state. AtlasVLA features a dual-memory architecture: a 4D Persistent World State Memory that lifts transient 2D observations into a globally updated, voxel-hashed spatial state to resolve visual blind spots, and an Ego-Working State Memory that tracks historical ego state and task progress. By conditioning a diffusion transformer (DiT) on this joint World-Ego state, AtlasVLA enables robust spatial reasoning. Extensive evaluations across LIBERO, RLBench, and real-world benchmarks demonstrate that AtlasVLA achieves state-of-the-art performance using solely a wrist camera. Remarkably, it decisively outperforms multi-view baselines, yielding absolute success rate improvements of 9.4% on LIBERO-Long and 17.5% in real-world long-horizon tasks.