AI 中文总结
研究提出FabriVLA模型,结合视觉-语言主干与流匹配动作头,经单阶段联合优化训练。在Meta-World MT50基准测试中,该模型基于1B规模VLM,不依赖数十亿参数主干,实现90.0%平均成功率,展现强大性能。
AI 中文摘要
我们提出了FabriVLA,一种用于精确多任务操作的轻量级视觉-语言-动作模型。FabriVLA将InternVL3.5视觉-语言主干与流匹配动作头相结合,该动作头具有跨动作令牌的门控自注意力和浅层VLM层融合以丰富空间上下文。该模型通过从预训练的VLM和随机初始化的动作头进行单阶段联合优化来训练。在跨越50个不同操作任务的Meta-World MT50基准测试中,FabriVLA实现了90.0%的平均成功率,表明基于1B规模VLM构建的紧凑VLA无需依赖数十亿参数的VLA主干即可实现强大性能。
英文摘要
Vision-Language-Action (VLA) models have become a leading paradigm for general purpose robotic manipulation, but their computational cost and limited uncertainty awareness hinder practical deployment. We present FabriVLA, a lightweight VLA that fuses shallow and intermediate VLM layers to preserve fine-grained visual features, and gates self-attention among action tokens so that its flow matching head admits inter step structure only as far as training warrants. Trained end-to-end in a single stage, FabriVLA reaches a state-of-the-art 90.0\% average success on Meta-World MT50 with only 0.88B parameters. We further introduce Joint Conformal Action Chunk Calibration (JCAC), a post-training method that augments a frozen policy with a lightweight residual scale head. From a single policy query, JCAC turns a learned elementwise error scale into a set that covers the whole executed action prefix at a user chosen confidence level, 3.3$\times$ tighter in mean radius than an unconditional conformal set. On LIBERO-Safety, these bounds rank rollouts by risk before execution, supporting risk ranked review. Together, FabriVLA and JCAC provide a lightweight and auditable framework for multi-task manipulation with calibrated action uncertainty.