arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.02958cs.ROcs.AI

ValueFormer:面向半自主视觉-语言-动作策略的、具有阶段感知标签的因果变换价值函数

ValueFormer: A Causal Transformer Value Function with Stage-Aware Labels for Semi-Autonomous Vision-Language-Action Policies

Inkyu Sa, Konstantin Stulov, Rajat Bhageria

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出 ValueFormer,一种带阶段感知标签的因果变换价值函数,用于半自主 VLA 策略,在真实机器人双手机器人三明治组装任务上提升了任务完成率并降低了服务成本。

中文摘要 AI 辅助

通过行为克隆训练的视觉-语言-动作(VLA)策略会出现隐性失效:仅从动作流来看,失败的 rollout 看起来和正常推进的 rollout 差别不大,因为模仿学习本身不提供进度的概念。强化学习本可提供进度概念,但在此场景下不实用——真实机器人的训练经验成本高昂,且可变形食物难以模拟。廉价的替代方案是使用终端成功/失败标记,该标记原则上可学习,但过于稀疏,无法说明 rollout 何时出错。本文认为,困难在于每帧标签而非架构:要发挥作用,标签必须是密集、连续且形状正确的。我们提出 ValueFormer,这是一种紧凑的、与策略无关的因果变换,基于冻结的 DINOv3 主干网络,在一次前向传播中输出两个每帧信号:用于优势估计的平滑蒙特卡洛价值 V_mc,以及用于在线错误检测的尖锐二元价值,这两个目标在设计上方向相反。失败的 episode 被标记为具有阶段感知的“成功后衰减”回报,保留失败阶段前的成功曲线;错误检测则由错误区间而非单一失败时间进行监督,因此策略可恢复的错误也会携带信号。在真实机器人双手机器人三明治组装任务(1427 个 episode)上,基于评论者的每帧训练权重将任务完成率从 70%提升至 85%(n=20 时处于噪声范围内),而批量 bf16 编码器将实时服务成本降低 3~5 倍,使评论者可在单个 GPU 上与策略以 2 Hz 运行。

英文摘要

Vision-Language-Action (VLA) policies trained by behavior cloning fail silently: from the action stream alone, a collapsing rollout looks much like one making clean progress, because imitation supplies no notion of progress. Reinforcement learning would supply one, but it is impractical here, where real-robot experience is costly and deformable food resists simulation. The cheap alternative, a terminal success / failure bit, is learnable in principle yet far too sparse to say when a rollout went wrong. We argue that the per-frame label, not the architecture, is the hard part: to be useful it must be dense, continuous, and correctly shaped. We present ValueFormer, a compact policy-agnostic causal transformer over a frozen DINOv3 backbone that emits two per-frame signals in one forward pass: a smooth Monte Carlo value, V_mc, for advantage estimation and a sharp binary value for online mistake detection, targets that pull in opposite directions by design. Failed episodes are labeled with a stage-aware, success-then-decay return that preserves the success curve before the failure stage, and detection is supervised from mistake intervals rather than a single failure time, so mistakes the policy recovers from also carry signal. On a real-robot bimanual sandwich-assembly task 1,427 episodes), a critic-derived per-frame training weight lifts task completion from 70% to 85% (within noise at n=20), and a batched bf16 encoder cuts the live serving cost 3~5 times so the critic runs at 2 Hz alongside the policy on a single GPU.

补充信息

↑