发表机构
Foshan University; The Chinese University of Hong Kong, Shenzhen; DexForce; South China University of Technology(佛山大学; 香港中文大学(深圳); DexForce; 华南理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
TaskAnchor通过历史条件化视觉细化和里程碑监督的任务状态坐标,解决反应式VLA在长时程操作中的任务状态混叠问题,在RMBench上成功率提升至基线的4.9–5.5倍。
AI 中文摘要
反应式视觉-语言-动作(VLA)模型在长时程操作中面临困难,因为视觉上相似的观测可能对应不同的动作,这取决于任务阶段或交互历史。我们将这种歧义称为任务状态混叠,并引入TaskAnchor,一种轻量级适配器,将预训练的VLA锚定在执行历史中。TaskAnchor将历史条件化的视觉细化与里程碑监督的任务状态坐标(表示执行语义阶段的标量)相结合。这些信号分别通过原生视觉和语言接口注入,无需引入显式规划器或修改动作生成机制。在RMBench上,TaskAnchor的平均成功率约为已发布的π₀.₅和X-VLA基线的4.9–5.5倍,在RoboMemArena和真实机器人上也有持续提升。对于π₀.₅,每个动作块的额外延迟仅为2.08毫秒。
英文摘要
Reactive vision-language-action (VLA) policies suffer from task-state aliasing in long-horizon manipulation, where identical multimodal inputs call for distinct, context-dependent actions. Given that pretrained VLAs already possess rich control primitives to express diverse behaviors, we hypothesize that the execution bottleneck lies not in policy capacity, but in input ambiguity. In this paper, we propose TaskAnchor, a lightweight adapter that grounds task state by injecting execution context into the VLA's native input space. During post-training, TaskAnchor learns to represent the semantic execution stage as a milestone-supervised coordinate prepended to the language instruction, while incorporating fine-grained historical evidence via a residual update to the current visual tokens. This formulation avoids generating complex subtask instructions and leaves the backbone architecture unchanged. Across long-horizon benchmarks, TaskAnchor delivers substantial gains, achieving approximately 6 times the average success rate of the pi0.5 and X-VLA baselines on RMBench and more than doubling the task success rate of pi0.5 on RoboMemArena. Real-robot experiments further validate reliable multi-stage execution, with the same policy adapting its subsequent behaviors using earlier human interactions as in-context cues. Our project website is available at https://taskanchor.netlify.app/.