RoboBRIDGE:一种将策略桥接至鲁棒现实世界机器人智能体的模块化框架
RoboBRIDGE: A Modular Framework for Bridging Policies to Robust Real-World Robotic Agents
浏览论文内容
中文总结 AI 辅助
该研究针对视觉-语言-动作模型部署为机器人智能体的缺陷,提出 RoboBRIDGE 模块化框架,通过五大协同模块结合预训练 VLAs 构建鲁棒智能体,在多基准和现实场景中性能优于现有方案。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型作为机器人操纵的可扩展方法受到越来越多的关注。这些模型是有效的动作预测器,但将其部署为机器人智能体时暴露出关键缺陷:缺乏故障恢复机制、长期执行一致性不足,以及对观测、任务或 embodiment 变化的鲁棒性有限。现有解决方案通过模型重新训练或特定环境模块单独解决这些限制,但需要的是一种通用框架,能系统地将预训练的 VLA 转化为机器人智能体。我们提出 RoboBRIDGE,这是一种模块化框架,在五个协同模块(即 Monitor、Perceptor、Planner、Controller 和 Robot Interface)之上提供编排层,以由现成组件(包括预训练 VLAs)组合出鲁棒机器人智能体。Monitor 将快速故障检测与分层恢复相结合,在错误级联前进行纠正。当环境偏离当前计划时,Planner 触发重新规划,同时 Perceptor 异步更新场景理解,避免执行停滞。在 Controller 中,基础技能微调通过专用 LoRA 适配器将操纵分解为领域不变的基础,降低了使用 VLA 时对领域变化的敏感性。在 LIBERO、RoboCasa 和涵盖多个机器人平台及 VLA 主干的现实案例研究中,RoboBRIDGE 的性能始终优于独立策略和先前增强型 VLA 部署。这些结果表明,可靠的机器人智能性并非仅来自动作预测器的扩展,而是来自围绕它们的结构化编排。
英文摘要
Vision-Language-Action (VLA) models have attracted growing interest as a scalable approach to robotic manipulation. While these models are effective action predictors, deploying them as robotic agents exposes critical gaps: no mechanism for failure recovery, inconsistent execution over long horizons, and limited robustness to shifts in observations, tasks, or embodiments. Existing solutions address these limitations individually through model retraining or environment-specific modules, yet what is needed is a general framework that systematically transforms a pretrained VLA into a robotic agent. We present RoboBRIDGE, a modular framework that provides an orchestration layer over five coordinated modules, namely Monitor, Perceptor, Planner, Controller, and Robot Interface, to compose robust robotic agents from off-the-shelf components, including pretrained VLAs. The Monitor pairs rapid failure detection with hierarchical recovery to correct errors before they cascade. When the environment diverges from the current plan, the Planner triggers replanning while the Perceptor updates scene understanding asynchronously, avoiding execution stalls. Within the Controller, primitive skill fine-tuning factors manipulation into domain-invariant primitives with dedicated LoRA adapters, reducing sensitivity to domain shifts when a VLA is used. Across LIBERO, RoboCasa, and real-world case studies spanning multiple robot platforms and VLA backbones, RoboBRIDGE consistently outperforms both standalone policies and prior augmented VLA deployments. These results suggest that reliable robotic agency does not arise from scaling action predictors alone, but from structured orchestration around them.
发表机构
- Sungkyunkwan University(成均馆大学)
机构由 AI 辅助整理,请以论文原文为准。