发表机构
Tongji University; The Hong Kong Polytechnic University; University of Electronic Science and Technology of China(同济大学; 香港理工大学; 电子科技大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对VLA策略中固定动作视界无法适应任务阶段变化的问题,提出基于Flow Matching去噪轨迹几何的自适应动作分块方法GeoAAC,无需额外训练即可调整视界,在仿真和真实任务中显著提升成功率。
AI 中文摘要
动作分块广泛用于视觉-语言-动作(VLA)策略中的动作生成与执行,然而现有方法通常采用固定的动作视界。在展开(rollout)过程中,不同任务阶段可能对动作连续性、控制精度和闭环反馈有不同需求,使得固定视界无法适应变化的控制要求。我们提出GeoAAC,一种基于几何的自适应动作分块方法,用于基于流的VLA策略,该方法根据当前动作预测的可靠性调整动作视界。我们表明,Flow Matching去噪轨迹的几何特性提供了用于刻画预测可靠性的过程级信息,且跨动作前缀的几何变化与预测不确定性保持正相关。GeoAAC利用这种前缀级几何构造视界级几何剖面,并仅通过单次生成即可自适应地确定动作视界,无需额外训练。在LIBERO、LIBERO-Pro、RoboCasa365及真实世界操作任务上使用GR00T N1.5和π0.5进行的实验表明,与固定动作视界基线和现有自适应方法相比,GeoAAC持续改进,包括在仿真中最多提升8.7个百分点,并将真实世界平均成功率从53.3%提升至74.4%。
英文摘要
Action chunking is widely used for action generation and execution in Vision-Language-Action (VLA) policies, yet existing approaches commonly use a fixed action horizon. During a rollout, different task stages may require different levels of action continuity, control precision, and closed-loop feedback, making a fixed horizon unable to accommodate changing control requirements. We propose \textbf{GeoAAC}, a geometry-based adaptive action chunking method for flow-based VLA policies that adjusts the action horizon according to the reliability of the current action prediction. We show that the geometry of Flow Matching denoising trajectories provides process-level information for characterizing prediction reliability, with geometric variation across action prefixes remaining positively correlated with predictive uncertainty. GeoAAC uses this prefix-wise geometry to construct a horizon-wise geometric profile and adaptively determine the action horizon from a single generation without additional training. Experiments with GR00T N1.5 and π0.5 on LIBERO, LIBERO-Pro, RoboCasa365, and real-world manipulation tasks show consistent improvements over fixed-action-horizon baselines and existing adaptive methods, including up to 8.7 percentage points in simulation and an increase in average real-world success rate from 53.3\% to 74.4\%.
Comments9 pages, 6 figures. Submitted to the IEEE International Conference on Robotics and Automation (ICRA) 2027