发表机构
University of Michigan; XPENG Robotics(密歇根大学; 小鹏机器人)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出主转向子空间(PSS),利用解码器响应构建低维控制基,通过软演员-评论家在线自适应冻结生成式机器人策略,在多个任务上加速收敛、提升性能,并成功应用于类人VLA策略。
AI 中文摘要
生成式机器人策略提供了富有表现力的行为先验,但通过在线交互更新大型扩散或流匹配模型成本高昂。潜在空间强化学习通过控制初始采样噪声来避免更新预训练生成器,然而高维噪声对解码动作可能产生强烈的各向异性效应。我们引入了主转向子空间(PSS),这是一种前向查询接口,通过有限差分解码器响应构建一个固定的低维控制基。软演员-评论家算法控制主导响应方向,而正交补空间则在每次查询时独立地从高斯先验中重新采样。在三个使用扩散和流匹配策略的RoboMimic任务中,响应谱显示出显著的集中性。在五对匹配的任务-生成器组合中,训练曲线表明,与全潜在控制相比,PSS通常收敛更快,表现出更稳定的后期训练行为,同时整体最终性能更强。受控的Diffusion-Square消融实验进一步表明,主导响应方向优于相同维度的随机和最小响应子空间。我们进一步将PSS与一个冻结的、闭源的3B参数视觉-语言-动作(VLA)策略集成到一个类人学习系统中,该系统具有同步转换收集、重置时间优化和延迟感知的异步部署。在探索性的螺丝刀放置评估中,冻结VLA策略的成功率为2/10次试验,而SAC+PSS自适应后为6/10次。这些结果支持将解码器响应几何作为冻结生成式机器人策略在线自适应的实用基础。
英文摘要
Generative robot policies provide expressive behavior priors, but updating a large diffusion or flow-matching model through online interaction is costly. Latent-space reinforcement learning avoids updating the pretrained generator by controlling its initial sampling noise, yet high-dimensional noise can have strongly anisotropic effects on decoded actions. We introduce Principal Steering Subspaces (PSS), a forward-query interface that constructs a fixed low-dimensional control basis from finite-difference decoder responses. Soft Actor-Critic controls the leading response directions, while the orthogonal complement is independently resampled from the Gaussian prior at each query. On three RoboMimic tasks with diffusion and flow-matching policies, response spectra reveal substantial concentration. Across five matched task-generator pairs, the training curves indicate that PSS generally converges faster and exhibits more stable late-training behavior than full-latent control, while achieving stronger final performance overall. Controlled Diffusion-Square ablations further show that leading-response directions outperform random and least-responsive subspaces of equal dimension. We further integrate PSS with a frozen, closed-source 3B-parameter vision-language-action (VLA) policy in a humanoid learning system with synchronous transition collection, reset-time optimization, and latency-aware asynchronous deployment. In an exploratory screwdriver-placement evaluation, success is observed in 2/10 trials for the frozen VLA policy and 6/10 after SAC+PSS adaptation. These results support decoder-response geometry as a practical basis for online adaptation of frozen generative robot policies.
Comments8 pages, 9 figures