arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.04933cs.ROcs.AIcs.LG

DiVeR:用于VLA测试时扩展的决策关键验证器学习

DiVeR: Decision-Critical Verifier Learning for VLA Test-Time Scaling

Seongheon Park, Heecheol Kim, Shulin Tian, Lilika Makabe, Namiko Saito, Katsushi Ikeuchi, Sharon Li, Yasuyuki Matsushita

首次发表
浏览论文内容

中文总结 AI 辅助

DiVeR通过估计状态决策关键性并重加权验证器学习,在不增加标注和交互成本的情况下,提升VLA策略的测试时扩展性能,在多个基准和真实机器人上持续提高任务成功率。

中文摘要 AI 辅助

扩展机器人数据和模型容量已经改善了视觉-语言-动作(VLA)策略,但进一步的进展受到机器人数据高成本的制约。验证器引导的测试时扩展提供了一种高效的替代方案,通过在推理时采样多个动作候选并选择最有可能导致任务成功的一个。现有的基于分类的验证器从轨迹级结果中学习,但将所有访问过的状态视为同等重要,尽管它们对候选区分的价值在轨迹中可能有所不同。在许多状态下,合理的动作相似且提供的区分信号有限,而只有稀疏的决策关键状态允许有意义的、不同的动作,这些动作可能显著影响下游结果。为了解决这个问题,我们提出了DiVeR,它从采样动作表示的离散度来估计决策关键性。DiVeR然后利用这一信号重新加权验证器学习,使其更关注动作选择最为关键的状态,而无需步骤级标注或额外的环境交互。在LIBERO、RoboCasa以及Franka Research 3机器人上的真实世界实验中,DiVeR通过更有效的验证器引导动作选择持续提高了任务成功率,同时增加了可忽略的验证器推理开销。

英文摘要

Scaling robot data and model capacity has improved Vision-Language-Action (VLA) policies, but further progress is constrained by the high cost of robotic data. Verifier-guided test-time scaling offers an efficient alternative by sampling multiple action candidates and selecting the one most likely to lead to task success at inference time. Existing classification-based verifiers learn from trajectory-level outcomes but treat all visited states equally, even though their value for candidate discrimination can vary across a trajectory. At many states, plausible actions are similar and provide limited discrimination signal, while only a sparse set of decision-critical states admits meaningfully different actions that can substantially affect downstream outcomes. To address this, we propose DiVeR, which estimates decision criticality from the dispersion of sampled action representations. DiVeR then uses this signal to reweight verifier learning toward states where action selection is most consequential, without requiring step-level annotations or additional environment interaction. Across LIBERO, RoboCasa, and real-world experiments on a Franka Research 3 robot, DiVeR consistently improves task success through more effective verifier-guided action selection, while adding negligible verifier inference overhead.

↑