arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35303cs.CV

PIVOT:面向多轮VLM智能体的枢轴感知在线策略自蒸馏

PIVOT: Pivot-Aware On Policy Self Distillation for Multi-Turn VLM Agents

Jiazhou Zhou, Hu Zhou, Yucheng Chen, Jinyuan Qu, Ying-Cong Chen, Lei Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对多轮VLM智能体训练中GRPO的稀疏奖励问题,提出PIVOT框架,内化枢轴定位与状态恢复,消除环境回滚,在五个基准上超越现有方法,准确率达0.90-0.92。

中文摘要 AI 辅助

基于可验证奖励的强化学习(RLVR)结合组相对策略优化(GRPO)被广泛用于多轮视觉语言模型(VLM)智能体训练,但其在均匀失败时存在零梯度静默问题,且采用粗粒度的回合级信用分配。尽管在线策略蒸馏(OPD)和在线策略自蒸馏(OPSD)利用事后信息缓解了稀疏奖励问题,但其底层机制仍不明确。通过在五个多轮VLM智能体基准上进行受控反事实回滚探测,我们发现OPSD/OPD的性能提升主要归因于枢轴步骤处的物理状态回滚,枢轴步骤定义为在剩余步数预算不足时第一个不可恢复的动作。然而,物理状态回滚在计算上代价高昂,且在实际环境中不可行。为弥合这一差距,我们提出枢轴感知内化视觉在线策略训练(PIVOT),这是一种强化学习框架,将枢轴定位和状态恢复直接内化到token级参数更新中,从而在RL训练期间消除环境回滚,并在测试时无需额外技能提示。PIVOT在单一架构中统一了三个功能角色:一个失败分析器非侵入式地定位枢轴步骤,并从视觉轨迹拼贴和动作日志中诊断失败模式;一个分离的教师模型在此特权诊断上下文下对失败token重新评分;一个学生模型优化联合GRPO和置信度门控OPD目标。在测试时,教师和分析器分支均被移除。在认知网格谜题、3D具身控制与导航以及生成推理等五个多轮VLM智能体任务上评估,PIVOT在Qwen2.5-VL-3B上达到0.90的总体准确率(比SFT+GRPO基线高8%,比先前最先进方法高5%),并在Qwen3-VL-2B上扩展到0.92(比SFT+GRPO基线高12%)。

英文摘要

Reinforcement learning with verifiable rewards (RLVR) via Group-Relative Policy Optimization (GRPO) is widely used for multi-turn VLM agent training, yet it suffers from zero-gradient silence on uniform failures and coarse episode-level credit assignment. While On-Policy Distillation (OPD) and On-Policy Self-Distillation (OPSD) mitigate sparse rewards using hindsight information, their underlying mechanisms remain poorly understood. Through controlled counterfactual rollback probes across five multi-turn VLM agent benchmarks, we reveal that performance gains in OPSD/OPD are largely driven by physical state rollback at the pivot step, defined as the first unrecoverable action without remaining step budget. However, physical state rollbacks are computationally prohibitive and infeasible in real-world environments. To bridge this gap, we present Pivot-Aware Internalized Visual On-Policy Training (PIVOT), an RL framework that internalizes pivot localization and state restoration directly into token-level parameter updates, eliminating environment rollbacks during RL training and additional skill hints at test time. PIVOT unifies three functional roles within a single architecture: a failure Analyzer non-invasively localizes the pivot step and diagnoses failure modes from visual trajectory collages and action logs; a detached Teacher re-scores failed tokens under this privileged diagnostic context; and a Student optimizes joint GRPO and confidence-gated OPD objectives. At test time, both Teacher and Analyzer branches are stripped. Evaluated on five multi-turn VLM agent tasks across cognitive grid puzzles, 3D embodied control and navigation, and generative reasoning, PIVOT achieves 0.90 overall accuracy on Qwen2.5-VL-3B (+8% over SFT+GRPO baseline and +5% over previous SOTA) and scales to 0.92 on Qwen3-VL-2B (+12% over SFT+GRPO baseline).

发表机构

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • International Digital Economy Academy (IDEA)(国际数字经济学院(IDEA))
  • The Hong Kong Polytechnic University(香港理工大学)
  • Nanyang Technological University(南洋理工大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑