arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14379cs.ROcs.AI

Reflex:面向反应关键型操作的快速且可预测的视觉-语言-动作模型

Reflex: Enabling Fast and Predictive Vision-Language-Action Models for Reaction-Critical Manipulation

  • Shanghai Jiao Tong University(上海交通大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuxuan Chen, Wanruo Zhang, Xiao Li

AI总结:

针对现有VLA模型忽略动态交互场景的问题,提出面向反应关键型操作的ReflexBench基准与无需大规模机器人数据预训练的ReflexVLA模型,其可提升动态操作性能并保持静态操作精度。

AI中文摘要:

视觉-语言-动作(VLA)模型近期在机器人操作中取得了令人瞩目的性能,但现有基准主要评估静态操作任务的泛化能力,且在很大程度上忽略了动态交互场景。为解决这一差距,我们提出了ReflexBench——一个面向反应关键型操作的基准,包含6项动态任务,并引入了将模拟器步进与机器人控制解耦的评估框架,支持同步和异步推理下的可配置延迟。基于ReflexBench,我们提出了ReflexVLA,这是一款无需大规模机器人数据预训练、专为反应关键型操作设计的高效VLA模型。ReflexVLA通过在视觉骨干网络中进行潜在未来预测和多帧时间融合来增强时间推理能力,同时通过批量视觉编码和CUDA Graph重放降低部署延迟。实验表明,ReflexVLA在持续提升动态操作性能的同时,在标准静态操作基准上保持了有竞争力的精度,真实世界实验进一步验证了其在实际部署条件下的有效性。项目网站:this https URL

英文摘要:

Vision-Language-Action (VLA) models have recently achieved promising performance in robotic manipulation. However, existing benchmarks mainly evaluate generalization on static manipulation tasks and largely overlook dynamic interaction scenarios. To address this gap, we present ReflexBench, a benchmark for reaction-critical manipulation. ReflexBench contains six dynamic tasks and introduces an evaluation framework that decouples simulator stepping from robot control while supporting configurable latency under synchronous and asynchronous inference. Building upon ReflexBench, we propose ReflexVLA, an efficient VLA model designed for reaction-critical manipulation without large-scale robot-data pretraining. ReflexVLA enhances temporal reasoning through latent future prediction and multi-frame temporal fusion within the vision backbone, while reducing deployment latency through batched visual encoding and CUDA Graph replay. Experiments show that ReflexVLA consistently improves dynamic manipulation performance while maintaining competitive accuracy on standard static manipulation benchmarks, and real-world experiments further demonstrate its effectiveness under practical deployment conditions. Project website: https://reflexvla.github.io

补充信息

相关深度报道

↑