arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.12287cs.RO

减少时间冗余以实现高效视觉-语言-动作推理

Reducing Temporal Redundancy for Efficient Vision-Language-Action Inference

Yuzhou Wu, Yuxin Zheng, Muchun Niu, Yishan Yang, Tianhao Liu, hanwen kang, Jiajian Jing, Linfeng Zhang, Chuan Wen

首次发表
浏览论文内容

中文总结 AI 辅助

研究针对VLA模型推理延迟高的问题,识别出时间冗余来源,提出系统级加速策略,在感知和动作生成上减少计算,经实验验证该策略能加速超2倍且保持高性能。

中文摘要 AI 辅助

视觉-语言-动作(VLA)模型在机器人操作中具有强大的泛化能力,但其高推理延迟限制了实时部署。我们识别出了现有VLA管道中时间冗余的两个主要来源:高度相似连续帧的重复视觉编码和基于扩散策略中的多步迭代采样。为解决此问题,我们提出一种系统级加速策略,减少感知和动作生成中的计算。在感知方面,仅增量更新对应动态场景区域的令牌而非重新编码整个帧。在策略方面,通过以效率为导向的训练将扩散采样压缩为紧凑的两步调度,同时保持动作精度。在多个平台上的实验表明,在保持高性能的同时加速超过2倍,在一般操作基准测试中成功率高达98%。

英文摘要

Vision-Language-Action (VLA) models exhibit strong generalization for robotic manipulation, yet their high inference latency limits real time deployment. We identify two primary sources of temporal redundancy in existing VLA pipelines: repeated visual encoding of highly similar consecutive frames and multi step iterative sampling in diffusion based policies. To address this, we propose a system level acceleration strategy that reduces computation in both perception and action generation. On the perception side, we incrementally update only tokens corresponding to dynamic scene regions instead of re-encoding entire frames. On the policy side, we compress diffusion sampling into a compact 2-step schedule through efficiency oriented training while preserving action precision. Experiments on Libero, RobotWin, and Real Robot Platforms demonstrate over 2 times speedup while maintaining high performance, achieving up to 98% success rate on general manipulation benchmarks. Our codes will be released on Github.

补充信息

↑