arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.08761cs.AIcs.RO

VeriFine:具身推理中自我改进的验证扩展

VeriFine: Scaling Verification for Self-Improvement in Embodied Reasoning

  • NVIDIA(英伟达)
  • UCLA(加州大学洛杉矶分校)
  • UC Berkeley(加州大学伯克利分校)
  • Stanford University(斯坦福大学)

机构由 AI 辅助整理,请以论文原文为准。

Zewei Zhou, Rachel Luo, Yulong Cao, Chaowei Xiao, Chensheng Peng, Boyi Li, Thomas Tian, Zheng Lian, Yan Wang, Jiaqi Ma, Boris Ivanovic, Marco Pavone, Wenhao Ding

AI总结:

VeriFine通过策略、课程与评判者共同进化扩展验证,在具身推理中实现持续自我改进,实验证明驾驶和导航任务中策略与评判能力均持续提升。

AI中文摘要:

自我改进的策略会不断暴露新的失败模式,从而改变其评判者必须能够验证的内容。然而,当前固定的评判者既限制了优化反馈,也限制了有用训练样本的发现,进而限制了进一步的自我改进。这一挑战在具身推理中尤为严峻,因为可靠的评估必须考虑空间基础、因果推理和安全性感知的决策。我们引入了VeriFine,一个智能体框架,通过策略、训练课程和评判者的共同进化来扩展验证。策略改进循环使用基于规则的评判者来诊断反复出现的失败,构建自适应课程,并优化策略。当进展停滞且验证成为瓶颈时,评判者改进循环会选择性地在信息丰富的失败案例上查询人类指导,并通过协同校准来完善评判者,其中人类和智能体解决分歧并趋同于物理推理的客观规则。修订后的评判者随后指导下一阶段的数据选择和策略优化。在驾驶和机器人导航任务上的实验表明,在强化学习和监督微调中,策略和评判者能力均持续自我改进。这些结果展示了随着策略失败模式的演变,扩展验证如何支持持续的自我改进。

英文摘要:

Self-improving policies continually expose new failure patterns, changing what their judges must be able to verify. However, current fixed judges constrain both optimization feedback and the discovery of useful training examples, limiting further self-improvement. This challenge is even more acute in embodied reasoning, where reliable evaluation must account for spatial grounding, causal reasoning, and safety-aware decision-making. We introduce VeriFine, an agent harness framework that scales verification through the co-evolution of the policy, training curriculum, and judge. The Policy Improvement Loop uses a rubric judge to diagnose recurring failures, construct an adaptive curriculum, and optimize the policy. When progress plateaus and verification becomes a bottleneck, the Judge Improvement Loop selectively queries human guidance on informative failure cases and refines the judge through coactive calibration, in which humans and agents resolve disagreements and converge toward the objective rubric of physical reasoning. The revised judge then guides the next stage of data selection and policy optimization. Experiments on driving and robot navigation tasks demonstrate continuous self-improvement in both policy and judge capability across reinforcement and supervised fine-tuning. These results show how scaling verification supports continuous self-improvement as policy failure patterns evolve.

补充信息

相关深度报道

↑