AI 中文总结
LIRA是面向VLA模型的局部跨层信息路由机制,通过深度感知路由VLM特征,在多机器人操作基准及真实场景中提升了动作预测性能与分布偏移下的鲁棒性。
AI 中文摘要
视觉-语言-动作(VLA)模型将预训练视觉-语言模型(VLM)的表征转换为机器人动作,但将中间VLM特征路由到动作解码器的接口仍未得到充分探索。现有设计要么仅暴露表征层级的狭窄部分,要么将每个解码器块与单个VLM层刚性匹配,限制了对跨深度互补任务证据的访问。我们提出LIRA,一种局部跨层动作条件机制,将VLM到动作的条件化表述为深度感知的信息路由。LIRA基于从中间VLM状态派生的任务标记特征和LIRA查询特征运行,为每个并行融合块分配以其对应VLM层为中心的深度对齐局部窗口。并行融合块聚合相邻LIRA查询特征,将其与任务标记特征及本体感受输入整合后进行动作预测。该路由接口保留了骨干架构、动作解码器和监督训练方案不变。在LIBERO、LIBERO-Plus、CALVIN ABC→D及真实世界操作任务中,LIRA在0.5B参数配置下,较VLA-Adapter基线提升了主要聚合指标;在LIBERO-Plus的零样本迁移任务中,LIRA将平均成功率从59.1%提升至78.0%,增幅达18.9个百分点,表明其在可控分布偏移下的鲁棒性得到提升。
英文摘要
Vision-Language-Action (VLA) models transform representations from pretrained vision-language models (VLMs) into robot actions, yet the interface that routes intermediate VLM features into action decoders remains underexplored. Existing designs either expose only a narrow part of the representation hierarchy or rigidly match each decoder block to one VLM layer, restricting access to complementary task evidence across depths. We introduce LIRA, a local cross-layer action-conditioning mechanism that formulates VLM-to-action conditioning as depth-aware information routing. LIRA operates on task-token features and LIRA Query features derived from intermediate VLM states, then assigns each Parallel Fusion Block a depth-aligned local window centered on its corresponding VLM layer. Parallel Fusion Blocks aggregate neighboring LIRA Query features and integrate them with task-token features and proprioceptive inputs before action prediction. This routing interface leaves the backbone architecture, action decoder, and supervised training recipe unchanged. Across LIBERO, LIBERO-Plus, CALVIN ABC$\rightarrow$D, and real-world manipulation, LIRA improves the principal aggregate metrics over the VLA-Adapter baseline under the same 0.5B-parameter configuration. In zero-shot transfer to LIBERO-Plus, LIRA increases average success from 59.1% to 78.0%, an 18.9-point gain indicating improved robustness under controlled distribution shifts.
Comments9 pages, 4 figures. Code and model checkpoints will be released upon acceptance of the paper