arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

三思而后行:将树搜索提炼为冻结视觉语言动作模型的动作评估

Look Before You Leap: Distilling Tree Search into Action Evaluation for Frozen VLA Models

Xinyi Xie, Zican Hu, Zhanyu Liu, Yicheng Dong, Wenhao Wu, Zhenhong Sun, Haoran Li, Chunlin Chen, Zhi Wang, Pichao Wang

arXiv 2607.03751首次发表:更新:

发表机构

Nanjing University; Australian National University; Institute of Automation, Chinese Academy of Sciences; Nvidia(南京大学; 澳大利亚国立大学; 中国科学院自动化研究所; 英伟达)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究发现视觉语言动作(VLA)模型泛化能力差的关键瓶颈在于动作评估。提出SVA框架,通过蒙特卡洛树搜索探索输出分布,将知识提炼到轻量级Q值模型,提升任务成功率,且保持泛化能力。

AI 中文摘要

视觉语言动作(VLA)模型通过大规模预训练获得广泛的具身能力,但其泛化能力比语言模型和视觉语言模型更脆弱。我们发现一个关键瓶颈:VLA失败不仅源于动作生成,还源于动作评估。受此启发,我们提出了SVA,一个简单的框架,为冻结的VLA策略配备长期后果意识。实验表明,SVA在保持VLA主干泛化能力的同时,显著提高了任务成功率。

英文摘要

Vision-Language-Action (VLA) models acquire broad embodied capabilities through large-scale pretraining, yet their generalization remains far more fragile than that of LLMs and VLMs. The prevailing remedy, post-training via supervised fine-tuning or reinforcement learning, improves task-specific performance but narrows the generalist capability that makes pretraining valuable. We identify a key bottleneck: VLA failures stem not only from action generation but also from action evaluation. A diagnostic pass@k study confirms that frozen VLAs already contain competent behaviors in their output distribution, with overall success rates rising from 33% at pass@1 to 92% at pass@32. Inspired by this, we propose SVA (Search, Value, and Act), a simple framework that equips frozen VLA policies with long-term consequence awareness. SVA first uses Monte-Carlo tree search in simulation to sample and explore the VLA's output distribution under a finite search budget, collecting diverse trajectories annotated with empirical returns; this knowledge is then distilled into a lightweight Q-value model that predicts the expected consequence of candidate actions; at deployment, the frozen VLA proposes multiple candidates and the evaluator selects the one with the highest uncertainty-regularized Q-value, requiring no simulator access. By decoupling action proposal from consequence evaluation, SVA preserves the generalization capacity of the VLA backbone while improving task success. Across embodied benchmarks, SVA consistently improves generalization on unseen tasks and exhibits strong test-time scaling; across real-robot manipulation tasks, it improves average success from 42.8% to 55.6% (+12.8 points). Strikingly, SVA enables a 9B VLA to outperform a 27B VLA by 7 points at 27% lower inference latency, suggesting that scaling test-time evaluation is more cost-effective than scaling model size.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑