AI 中文总结
本文提出环境感知模型选择(EMS)自适应VLA推理框架,通过解耦双系统与强化学习切换策略,在LIBERO基准及真实双臂任务中平衡了VLA的推理速度与任务成功率。
AI 中文摘要
具身智能既需要长程推理能力,又需要实时闭环响应能力。近期的双系统视觉-语言-动作(VLA)架构结合了快速反应控制与慢速 deliberative 推理,以平衡推理速度与任务成功率。然而,现有的双进程VLA将快速模块与慢速模块的中间表示紧密耦合,需要端到端联合训练,限制了模块化、可扩展性及灵活的系统切换。本文提出环境感知模型选择(EMS),一种自适应VLA推理框架,通过环境感知模型选择在两个完全解耦、不同规模的系统间切换。大规模 deliberative 系统提供全局一致的轨迹规划以确保任务成功,而轻量型反应系统实现高频闭环控制。基于强化学习的切换策略根据实时反馈动态选择调用哪个系统,实现慢速系统的稀疏使用,从而平衡预训练知识利用与运行时效率。与现有分层VLA框架相比,本文设计具有三个关键优势:(1)完全解耦的模块化双系统架构,支持即插即用的模型替换;(2)自适应的环境感知切换策略;(3)用于响应式闭环控制的高频推理。本文在仿真和真实环境中对EMS进行了广泛评估:在LIBERO基准测试中,EMS达到与大规模基线相当的成功率,同时将有效动作频率提升至93.4 Hz;该框架在真实世界的双臂操作任务中进一步展现出强大的可扩展性,在保持稳健性能的同时缩短了任务完成时间。
英文摘要
Embodied intelligence demands both long-horizon reasoning and real-time closed-loop responsiveness. Recent dual-system Vision-Language-Action (VLA) architectures combine fast reactive control with slow deliberative reasoning to balance inference speed and task success rate. However, existing dual-process VLAs tightly couple the fast module to intermediate representations of the slow module, necessitating end-to-end joint training and limiting modularity, extensibility and flexible system switching. In this paper, we propose Environment-aware Model Selection (EMS), an adaptive VLA inference framework that switches between two fully decoupled systems of different scales through environment-aware model selection. The large-scale deliberative system provides globally consistent trajectory planning to ensure task success, while a lightweight reactive system enables high-frequency closed-loop control. A reinforcement-learning-based switching policy dynamically selects which system to invoke based on real-time feedback, enabling sparse use of the slow system and thereby balancing pretrained knowledge utilisation with runtime efficiency. Our design offers three key advantages over prior hierarchical VLA frameworks: (1) a fully decoupled and modular dual-system architecture that supports plug-and-play model replacement; (2) an adaptive, environment-aware switching strategy; (3) high-frequency inference for responsive closed-loop control. We extensively evaluate EMS in both simulation and real-world environments. On the LIBERO benchmark, EMS achieves success rates comparable to the large-scale baseline while increasing the effective action frequency to 93.4 Hz. The framework further demonstrates strong extensibility in real-world dual-arm manipulation tasks, where it accelerates task completion while maintaining robust performance.