发表机构
KAIST; SKKU; GIST(韩国科学技术院; 成均馆大学; 光州科学技术院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出伴随引导流(AGF),将轨迹感知的评论家引导摊销为轻量级网络,在保持预训练VLA策略的同时,通过最优控制伴随状态实现高效引导,在多个基准上提升性能并显著降低计算开销。
AI 中文摘要
基于流的视觉-语言-动作(VLA)策略通常通过行为克隆进行训练,因此不会显式优化长期任务回报。评论家引导将生成过程导向更高价值的动作,但现有方法通过采样器的一步代理对评论家进行微分,并在每个流步骤中反向传播评论家集成。相比之下,本文提出伴随引导流(AGF),该方法将轨迹感知的评论家引导摊销到一个轻量级引导网络中,同时保留预训练的VLA策略。具体而言,我们将评论家引导的流生成建模为一个确定性最优控制问题,其最优引导是一个伴随状态,该伴随状态将终端评论家梯度通过剩余流反向传播,并在保持VLA和评论家均冻结的同时,将引导网络回归到该伴随状态上。这种设计在训练期间提供了有利的内存和吞吐量扩展,推理时每步只需一次引导网络前向传播,无需评论家集成、反向传播或伴随计算。在LIBERO、RoboCasa和LIBERO-Pro上,AGF一致地改进了预训练的VLA,与评论家引导和策略微调基线相比保持竞争力,并且在跨任务部署单一引导强度时是最稳健的方法。与QGF相比,AGF每引导步骤运行快3.6倍,参数少7.0倍,性能相当甚至更好,表明评论家引导可以是轨迹感知且轻量级的。
英文摘要
Flow-based Vision-Language-Action (VLA) policies are typically trained by behavior cloning and thus do not explicitly optimize long-term task return. Critic guidance steers generation toward higher-value actions, but existing methods differentiate the critic through a one-step surrogate of the sampler and back-propagate a critic ensemble at every flow step. In contrast, here we propose Adjoint Guidance Flow (AGF), which amortizes trajectory-aware critic guidance into a lightweight guidance network while preserving the pretrained VLA policy. Specifically, we formulate critic-guided flow generation as a deterministic optimal control problem, whose optimal guidance is a costate that carries the terminal critic gradient back through the remaining flow, and regress the guidance network onto this costate while keeping both the VLA and critic frozen. This design provides favorable memory and throughput scaling during training, and inference needs one guidance-network forward pass per step, without the critic ensemble, back-propagation, or adjoint computation. Across LIBERO, RoboCasa, and LIBERO-Pro, AGF consistently improves pretrained VLAs, remains competitive with critic-guidance and policy-fine-tuning baselines, and is the most robust method when a single guidance strength is deployed across tasks. Compared with QGF, AGF runs $3.6\times$ faster per guidance step with $7.0\times$ fewer parameters, with comparable and even better performance, showing that critic guidance can be trajectory-aware and lightweight.