arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

世界校准的提议到行动流程用于视觉-语言-行动模型

World-Calibrated Proposal-to-Action Flow for Vision-Language-Action Models

Jie He, Wei Li, Junwen Tong, Rui Shao, Wei-Shi Zheng, Liqiang Nie

arXiv 2610.02323首次发表:更新:

发表机构

Harbin Institute of Technology, Shenzhen; State Key Laboratory of Mobile Network and Mobile Multimedia Technology, ZTE, China; Sun Yat-sen University(哈尔滨工业大学(深圳); 中兴通讯移动网络和移动多媒体技术国家重点实验室; 中山大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出ProAct框架,通过世界校准的提议到行动流程,使生成源可预测,保持运动连续性并校准各向异性源,提升VLA模型性能并减少计算开销。

AI 中文摘要

基于流的视觉-语言-行动(VLA)策略通过从任务无关的各向同性高斯源传输样本来生成行动块。由于该源既不依赖于最近的执行情况,也不依赖于预测的未来演化,(i)它丢弃了最近执行的运动所建立的局部连续性。(ii)即使引入了预测性世界表示,它们通常也只调节传输动力学,而不决定生成的起点、允许偏离的程度或沿哪些行动方向扩展。基于这一观察,我们提出了ProAct,一个世界校准的提议到行动框架,使生成源本身变得可预测。(i)为了保持运动连续性,一个轻量级提议专家通过一步运动锚定的端点流匹配将最近的动作转换为场景感知假设,在演示行动流形附近初始化生成。(ii)为了联合捕捉预期的场景演化和提议-未来兼容性,一个前瞻性世界专家将假设视为软运动先验,同时预测与任务一致的潜在未来。(iii)从这种兼容性中,模型校准一个以提议为中心的各向异性源,其中受限的每步范围控制允许的偏差,而迹归一化的低秩几何在条件数预算下分配对耦合的平移、旋转和夹爪方向的细化。与π0.5相比,ProAct在模拟和现实世界任务中提高了性能,同时将去噪步骤减少50%,推理延迟降低最多25.8%,吞吐量提高最多34.8%。

英文摘要

Flow-based Vision-Language-Action (VLA) policies generate action chunks by transporting samples from a task-agnostic isotropic Gaussian source. As this source is conditioned on neither recent execution nor predicted future evolution, (i) it discards the local continuity established by recently executed motion. (ii) Even when predictive world representations are introduced, they often only condition the transport dynamics rather than determine where generation starts, how far it may deviate, or along which action directions it may expand. Building on this observation, we introduce ProAct, a world-calibrated proposal-to-action framework that makes the generative source itself predictable. (i) To preserve motion continuity, a lightweight Proposal Expert converts recent actions into a scene-aware hypothesis via one motion-anchored endpoint flow-matching step, initializing generation near the demonstrated action manifold. (ii) To jointly capture intended scene evolution and proposal-future compatibility, a prospective World Expert treats the hypothesis as a soft motion prior while predicting the task-consistent latent future. (iii) From this compatibility, the model calibrates a proposal-centered anisotropic source, where a bounded per-step extent controls the allowed deviation and a trace-normalized low-rank geometry under a condition-number budget allocates refinement over coupled translation, rotation, and gripper directions. Compared with $π_{0.5}$, ProAct improves performance across simulation and real-world tasks while reducing denoising steps by 50%, inference latency by up to 25.8%, and increasing throughput by up to 34.8%.

Comments25 pages, 10 figures. Project page: https://github.com/JiuTian-VL/ProAct-page

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑