发表机构
Li Auto Inc.(理想汽车)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对思维链推理会降低强具身智能体性能的问题,区分显式推理与隐式决策计算,提出先决策后解释的DTE范式,在多任务上优于基线,证明显式CoT更适用于训练阶段塑造决策而非推理阶段介导动作。
AI 中文摘要
思维链(Chain-of-Thought, CoT)推理正越来越多地被集成到视觉-语言-动作(Vision-Language-Action, VLA)模型中,但它反而可能降低性能更强的具身智能体的表现。我们通过区分显式推理与隐式决策计算(即直接支撑动作预测的、基于感知的计算),研究了这种依赖智能体能力的效应。在标准的“先思考后行动”(think-then-act, TTA)范式下,固定视觉输入与模型参数的同时干预生成的CoT会导致性能大幅崩溃,导航F1值从72.14%降至11.84%,证明了显式推理对动作生成的强烈影响。随后我们提出“先决策后解释”(decide-then-explain, DTE)方法,在生成解释之前先预测动作,并引入视觉条件贡献(Visual Conditional Contribution, VCC)与推理条件贡献(Reasoning Conditional Contribution, RCC)来刻画由此产生的决策过程。在自动驾驶与机器人操作任务中,DTE的表现始终优于TTA和传统的无CoT基线,同时表现出对基于感知的计算更强的依赖性。进一步的“TTA训练、DTE推理”实验表明,这种优势并非仅来源于新分解方式下的重新训练。我们的结果显示,对于性能较强的具身智能体,显式CoT可能更适合用于在训练阶段塑造决策计算,而非在推理阶段介导动作生成。代码链接:https://github.com/ocean-luna/openvla-decide-then-explain。
英文摘要
Chain-of-Thought (CoT) reasoning is increasingly incorporated into Vision-Language-Action (VLA) models, yet it can degrade the performance of stronger embodied agents. We investigate this capability-dependent effect by distinguishing explicit reasoning from latent decision computation, i.e., perception-grounded computation that directly supports action prediction. Under the standard \textit{think-then-act} (TTA) paradigm, intervening on the generated CoT while fixing the visual input and model parameters causes a substantial performance collapse, with Navigation F1 dropping from 72.14% to 11.84%, demonstrating the strong influence of explicit reasoning on action generation. We then propose \textit{decide-then-explain} (DTE), which predicts actions before generating explanations, and introduce Visual Conditional Contribution (VCC) and Reasoning Conditional Contribution (RCC) to characterize the resulting decision process. Across autonomous driving and robotic manipulation, DTE consistently outperforms TTA and conventional \textit{no-CoT} baselines, while exhibiting greater reliance on perception-grounded computation. Further TTA-trained, DTE-inference experiments show that this benefit is not solely attributable to retraining under the new factorization. Our results suggest that for capable embodied agents, explicit CoT may be better used to shape decision computation during training rather than mediate action generation at inference time. Code: https://github.com/ocean-luna/openvla-decide-then-explain.