arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39198cs.RO

DSDyn-VLA:具有运动感知、未来感知与实时校正的双流动态操作框架

DSDyn-VLA: A Dual-Stream Dynamic Manipulation Framework with Motion Perception, Future Awareness, and Realtime Correction

Wenhao Li, Xiu Su, Yu Han, Yichao Cao, Shan You, Chang Xu

首次发表
浏览论文内容

中文总结 AI 辅助

DSDyn-VLA提出慢-快双流动态操作框架,通过光流感知与未来状态机制解决VLA在动态场景中的感知、延迟和控制差距,并引入DynBench基准,显著降低失败率并提升成功率。

中文摘要 AI 辅助

尽管视觉-语言-动作(VLA)模型在静态任务中表现出色,但在物体处于运动状态的动态环境中(例如传送带操作)却难以胜任。我们识别出当前VLA模型在这些场景中存在的三个根本性局限:感知差距,即静态视觉输入缺乏时间运动线索;延迟差距,即推理延迟导致动作过时;以及控制差距,这是由于开环动作块执行缺乏实时调整所致。在这项工作中,我们提出了DSDyn-VLA,一个慢-快双流动态操作框架,它集成了运动感知的前瞻性规划与实时残差校正。慢速的Flow-Planner作为宏观规划器,通过用光流增强VLA以实现时间感知,并采用未来状态感知机制来预先抵消推理延迟,从而生成全局一致、运动感知的动作块。与此互补,快速的Res-Refiner采用轻量级强化学习策略,基于实时观测向规划的动作块注入高频闭环校正。此外,我们引入了DynBench,一个基于MuJoCo的动态物体操作基准,包含九个任务。大量实验表明,在Kinetix动态基准的高延迟设置下,DSDyn-VLA相比当前最先进方法将失败率降低了超过76%,同时在真实世界动态设置中实现了约为PI0.5成功率6倍的表现,在DynBench上约为5倍。我们将开源所有代码和权重。

英文摘要

While Vision-Language-Action (VLA) models excel in static tasks, they struggle in dynamic environments where objects are in motion (e.g., conveyor belt manipulation). We identify three fundamental limitations hindering current VLAs in these scenarios: the \textbf{perception gap}, where static visual inputs lack temporal motion cues; the \textbf{latency gap}, where inference delays render actions obsolete; and the \textbf{control gap}, caused by the open-loop action chunk execution without real-time adjustment. In this work, we propose \textbf{DSDyn-VLA}, a Slow-Fast \textbf{D}ual-\textbf{S}tream \textbf{Dyn}amic manipulation framework that integrates motion-aware foresighted planning with real-time residual correction. The slow \textbf{Flow-Planner} serves as a macro-planner. By enhancing the VLA with optical flow for temporal perception and a future state awareness mechanism to preemptively offset inference latency, it produces globally consistent, motion-aware action chunks. Complementing this, the fast \textbf{Res-Refiner} employs a lightweight RL policy to inject high-frequency, closed-loop corrections into the planned action chunks based on real-time observations. In addition, we introduce \textbf{DynBench}, a MuJoCo-based benchmark for dynamic object manipulation that comprises nine tasks. Extensive experiments demonstrate that DSDyn-VLA reduces the failure rate by over 76\% compared to current SOTA method in high-latency setting on the Kinetix dynamic benchmark, while achieving about 6$\times$ the success rate of PI0.5 in real-world dynamic settings and about 5$\times$ on DynBench. We will open-source all the code and weights.

发表机构

  • University of Sydney(悉尼大学)
  • Central South University(中南大学)
  • University of California, San Diego(加利福尼亚大学圣迭戈分校)
  • Ace Robotics

机构由 AI 辅助整理,请以论文原文为准。

↑