发表机构
The University of Hong Kong; ACE Robotics(香港大学; ACE机器人公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出StreamPI框架,为单帧VLA模型赋予时序推理能力,通过指令锚定时序建模、随机间隔流式训练等方法,在真实机器人与仿真基准任务上性能优于pi0.5。
AI 中文摘要
视觉-语言-动作(VLA)模型在机器人操纵任务中已展现出有效性,但pi0.5等当前最优模型采用单帧范式,限制了其保留过往观测及构建精准空间感知的能力。本文提出StreamPI,一种流式多模态时序建模框架,为单帧VLA模型赋予时序推理能力且不引入额外参数。其核心设计之一是指令锚定时序建模:将每个(视觉观测,语言指令)对视为原子时序单元,单元内的双向注意力实现跨模态融合,单元间的因果注意力保持自回归流式推理,确保语言指令在整个任务执行过程中作为持续语义锚点。为弥合同步训练与异步真实机器人部署的差距,本文提出随机间隔流式训练策略:合适的帧间间隔(例如每3帧)可实现更快速、流畅的动作执行,随机化间隔进一步提升对帧时序扰动的鲁棒性,支持实际异步部署。此外,借助LLM主干的长度外推能力,StreamPI可无缝继承预训练单帧权重,支持灵活的单帧与多帧推理。在涵盖记忆依赖、精准感知场景的真实机器人任务及仿真基准LIBERO上的实验表明,StreamPI在各类任务中均优于pi0.5。
英文摘要
Vision-Language-Action (VLA) models have demonstrated effectiveness in robot manipulation, yet state-of-the-art models such as pi0.5 operate under a single-frame paradigm, limiting their ability to retain past observations and develop precise spatial perception. In this paper, we propose StreamPI, a streaming multimodal temporal modeling framework that equips single-frame VLA with temporal reasoning capability without introducing any additional parameters. One core design is instruction-anchored temporal modeling. It treats each (visual observation, language instruction) pair as an atomic temporal unit: bidirectional attention within each pair enables cross-modal fusion, while causal attention across pairs preserves autoregressive streaming inference. This ensures the language instruction serves as a persistent semantic anchor throughout task execution. To bridge the gap between synchronous training and asynchronous real-robot deployment, we introduce a andom-interval streaming training strategy: a proper inter-frame interval (e.g., every 3 frames) enables faster and smoother action execution. Beyond this, randomizing the interval further improves robustness to frame-timing perturbations, supporting asynchronous deployment in practice. Furthermore, by leveraging the length extrapolation capability of the LLM backbone, StreamPI seamlessly inherits pretrained single-frame weights and supports flexible single-frame and multi-frame inference. Experiments on real-robot tasks spanning memory-dependent and precise perception scenarios, as well as the simulation benchmark LIBERO, demonstrate that StreamPI outperforms pi0.5 across diverse tasks.