发表机构
Shanghai Jiao Tong University; PrimeBot; Zhejiang University; ShanghaiTech University; National University of Singapore(上海交通大学; PrimeBot; 浙江大学; 上海科技大学; 新加坡国立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
REACT是一种滚动去噪与双解耦框架,用于优化基于VLA模型的反应式机器人控制,在RoboTwin 2.0仿真及真实任务中提升了任务成功率、降低延迟并生成更平滑轨迹。
AI 中文摘要
基于流的视觉-语言-动作(VLA)模型会生成动作块以实现时间连贯的机器人运动,但分块控制存在根本的闭环权衡:长动作块可提供平滑执行效果,而频繁重规划虽能提升反应性,却会导致动作不连续。本文提出REACT,一种滚动去噪框架,可在保留长程上下文的同时提升基于流的VLA模型的反应性。REACT无需从头重新生成整个动作块,而是维护一个具有交错流时间步的持久动作缓冲区。在每个控制步骤,利用最新观测值对整个时域进行去噪,执行最干净的动作块,将部分优化的未来块向前移位,并在尾部添加新噪声。因此,每个执行的动作块在部署前会经过多个近期观测值的优化。为支持实时控制,本文进一步提出双解耦技术,将感知、视觉语言模型(VLM)编码、DiT去噪与动作执行分离,在实际计算约束下实现高频观测更新与动作流传输。在RoboTwin 2.0仿真基准及涵盖双臂操作、多机器人平台动态控制的真实任务中,REACT相比频繁重规划与异步基线方法,提升了任务成功率、降低了反应延迟,同时生成更平滑的轨迹。
英文摘要
Flow-based vision-language-action (VLA) models generate action chunks for temporally coherent robot motion, but chunked control creates a fundamental closed-loop trade-off: long chunks provide smooth execution, whereas frequent replanning improves reactivity at the cost of action discontinuities. We introduce REACT, a rolling-denoising framework that makes flow-based VLAs more reactive while preserving long-horizon context. Instead of regenerating entire action chunks from scratch, REACT maintains a persistent action buffer with staggered flow timesteps. At each control step, the full horizon is denoised using the latest observation, the cleanest action block is executed, partially refined future blocks are shifted forward, and fresh noise is appended to the tail. As a result, each executed action block is refined across multiple recent observations before deployment. To support real-time control, we further introduce dual decoupling, which separates sensing, VLM encoding, DiT denoising, and action execution, enabling high-frequency observation updates and action streaming under practical compute constraints. Across the RoboTwin 2.0 simulation benchmark and real-world tasks spanning bimanual manipulation and dynamic control on multiple robot platforms, REACT improves task success and reduces reaction latency while producing smoother trajectories than frequent-replanning and asynchronous baselines.
CommentsAccepted to CoRL 2026 as Spotlight. Project page: https://react-vla.github.io