CereVLA:小脑启发的后果感知残差治理用于高效视觉-语言-动作执行
CereVLA: Cerebellum-Inspired Consequence-Aware Residual Governance for Efficient Vision-Language-Action Execution
浏览论文内容
中文总结 AI 辅助
CereVLA提出小脑启发的后果感知残差治理框架,通过轻量级残差细化与预测评估,在冻结VLA上抑制不利纠正,提升成功率并减少控制步数。
中文摘要 AI 辅助
动作分块的视觉-语言-动作(VLA)策略提高了推理效率,但在已提交的动作分块内有限的反馈可能导致累积的执行误差。残差适应可以在不重新训练VLA的情况下纠正此类偏差;然而,现有的纠正通常针对参考动作一致性进行优化,而未明确考虑其下游后果。为解决这一局限,我们提出了小脑启发的后果感知残差治理(CereVLA),一个将轻量级残差细化与预测性后果评估集成到冻结的VLA执行中的统一框架。首先通过基于流的残差细化生成纠正动作,然后通过循环状态空间模型和历史感知分类器评估其短期和区间视野后果。预测为不利的残差纠正被一个轻量级治理器选择性地抑制。与LIBERO-10和LIBERO-GOAL上最先进方法的比较证明了CereVLA的有效性。在SO-101上,相对于冻结的SmolVLA基线,CereVLA将任务成功率从57.5%提高到90.0%,并在成功试验中平均控制步数减少了19.6%。
英文摘要
Action-chunked vision-language-action (VLA) policies improve inference efficiency, but limited feedback within committed action chunks can lead to accumulated execution errors. Residual adaptation can correct such deviations without retraining the VLA; however, existing corrections are typically optimized for reference-action consistency without explicitly considering their downstream consequences. To address this limitation, we present Cerebellum-Inspired Consequence-Aware Residual Governance (CereVLA), a unified framework that integrates lightweight residual refinement and predictive consequence evaluation into frozen VLA execution. Corrective actions are first generated by flow-based residual refinement, and their short- and interval-horizon consequences are then evaluated by a recurrent state-space model and a history-aware classifier. Residual corrections predicted to be unfavorable are selectively suppressed by a lightweight governor. Comparisons with state-of-the-art methods on LIBERO-10 and LIBERO-GOAL demonstrate the effectiveness of CereVLA. On SO-101, CereVLA increases task success from 57.5% to 90.0% and reduces mean control steps by 19.6% among successful trials, relative to the frozen SmolVLA baseline.
发表机构
- Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
机构由 AI 辅助整理,请以论文原文为准。