arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当推理有助于行动:视觉-语言-动作策略中思维链的监控与引导

When Reasoning Helps Action: Monitoring and Steering Chain-of-Thought in Vision-Language-Action Policies

Sathwik Karnik, Joseph JR. Lee, Aryaman Gupta, Somil Bansal

arXiv 2610.00601首次发表:更新:

发表机构

Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出TRUST模型,通过监控和引导VLA策略的思维链推理,提升推理正确性并降低碰撞率,但发现推理改进不必然带来具身性能提升。

AI 中文摘要

具备推理能力的视觉-语言-动作(VLA)策略会暴露思维链(CoT)轨迹,这些轨迹看似解释并引导其动作,从而为通过推理监控与修正实现运行时安全提供了潜在接口。在本工作中,我们定义并操作化两个评估轴,以评估该接口何时能改善具身行为:可修正性(correctability),衡量不可靠推理在生成过程中能否被检测并改进;以及可行动性(actionability),衡量推理修正是否在预期方向上产生行为上有意义的改变。为实现可修正性,我们引入了面向效用引导思维链的令牌级奖励(TRUST),这是一个离线训练的价值模型,能够从部分前缀预测最终推理的正确性,并利用这些估计来监控和选择性引导冻结VLA策略中的推理生成。在Alpamayo 1.5驾驶VLA上,TRUST以88.9%的准确率监控正确性,并将推理正确性从75.9%提升至90.0%。在AlpaSim中基于基线定义的具有挑战性的子集上,TRUST相对于未引导策略将碰撞率降低了30.4%,最大轨迹误差降低了11.5%,优于计算量匹配的Best-of-4基线。在DeepThinkVLA操作VLA上,TRUST将抓取状态声明的正确性从69.3%提升至90.2%,将动作选择声明的正确性从68.8%提升至85.9%,然而在LIBERO-Plus上的闭环任务性能基本保持不变。实证分析揭示了Alpamayo 1.5中意图一致的行为效应,但在DeepThinkVLA中效应有限,这有助于解释这些不同的任务级结果。综合来看,我们的结果表明,推理正确性的提升并不自动意味着具身性能的提升,这促使在使用CoT作为运行时安全接口时,需要评估可修正性和可行动性。

英文摘要

Reasoning-enabled VLA policies expose chain-of-thought (CoT) traces that appear to explain and guide their actions, creating a potential interface for runtime safety through reasoning monitoring and correction. In this work, we define and operationalize two evaluation axes for assessing when this interface can improve embodied behavior: correctability, which measures whether unreliable reasoning can be detected and improved during generation, and actionability, which measures whether reasoning corrections produce behaviorally meaningful changes in the intended direction. To enable correctability, we introduce Token-level Reward for Utility-Steered Chain-of-Thought (TRUST), an offline-trained value model that predicts eventual reasoning correctness from partial prefixes and uses these estimates to monitor and selectively steer reasoning generation in frozen VLA policies. On the Alpamayo 1.5 driving VLA, TRUST monitors correctness with 88.9% accuracy and improves reasoning correctness from 75.9% to 90.0%. On a baseline-defined challenging subset in AlpaSim, TRUST reduces collision rate by 30.4% and maximum trajectory error by 11.5% relative to the unsteered policy, outperforming a compute-matched Best-of-4 baseline. On the DeepThinkVLA manipulation VLA, TRUST improves the correctness of grasp-state claims from 69.3% to 90.2% and action-choice claims from 68.8% to 85.9%, yet closed-loop task performance on LIBERO-Plus remains largely unchanged. Empirical analysis reveals intent-consistent behavioral effects in Alpamayo 1.5 but limited effects in DeepThinkVLA, helping interpret these different task-level outcomes. Together, our results show that gains in reasoning correctness do not automatically imply gains in embodied performance, motivating evaluation of correctability and actionability when using CoT as a runtime safety interface.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑