AI 中文总结
该研究针对现有VLA模型用CoT会降低控制性能的问题,提出通过上下文后训练和智能体工具使用赋予VLA语言能力的方法,在多模拟和真实机器人任务上实现SOTA性能与效率。
AI 中文摘要
视觉-语言-动作(VLA)模型已成为通用操纵任务的主流方案,但它们几乎普遍通过行为克隆进行训练:策略模仿以静态图像和固定指令为条件的专家动作块。一种自然的解决方案是通过文本思维链(CoT)注入显式推理。我们通过实验和分析表明,自由形式的文本思维链会降低低级控制性能:其产生的推理缺乏依据,延迟会破坏闭环时序,且关键的是,推理和动作令牌是针对冲突目标优化的,导致策略学会叙述而非执行动作。我们认为VLA所需的不是生成语言的能力,而是使用有依据的语言的能力。为此,我们引入了我们的方法(\textbf{\textbf{\textbf{ourmethod}}}),该框架通过以下方式为VLA赋予语言能力:(i)上下文后训练,其中感知证据被注入为结构化上下文,模型仅在动作上进行监督;(ii)智能体工具使用接口,其中策略查询开放词汇检测器、单目深度估计器和视觉-语言模型以主动获取任务相关信息。我们的数据引擎生成多样的、意译的且以证据为条件的空间描述,而非单一模板化标题,从而使策略学会解释其从未逐字见过的语言。在RoboCasa-GR1、SimplerEnv和LIBERO模拟基准,以及8个真实世界机器人操纵任务上,与匹配配置下基于CoT的方法相比,我们的方法在性能和效率上始终达到SOTA结果。
英文摘要
Vision-Language-Action (VLA) models have become the dominant recipe for generalist manipulation, yet they are almost universally trained by behavior cloning: a policy imitates expert action chunks conditioned on a static image and a fixed instruction. A natural remedy is to inject explicit reasoning through textual chain-of-thought (CoT). We show, both empirically and analytically, that free-form textual CoT degrades low-level control: the reasoning it produces is ungrounded, its latency breaks closed-loop timing, and, crucially, the reasoning and action tokens are optimized against conflicting objectives so that the policy learns to narrate rather than to act. We argue that what a VLA needs is not the ability to generate language, but the ability to consume grounded language. To this end we introduce \textbf{\ourmethod{}}, a framework that endows a VLA with language competence through (i) in-context post-training, in which perceptual evidence is injected as structured context and the model is supervised only on actions, and (ii) an agentic tool-use interface, in which the policy queries open-vocabulary detectors, monocular depth, and a vision--language model to actively acquire task-relevant information. Rather than emitting a single templated caption, our data engine produces diverse, paraphrased, and evidence-conditioned spatial descriptions, so that the policy learns to interpret language it has never seen verbatim. Across the RoboCasa-GR1, SimplerEnv, and LIBERO simulation benchmarks, together with 8 real-world robot manipulation tasks, our method consistently achieves SOTA results in both performance and efficiency when compared with CoT-based approaches under matched configurations.