arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LayerRoute:面向视觉-语言-动作策略的动作条件混合层路由

LayerRoute: Action-Conditioned Mixture-of-Layers Routing for Vision-Language-Action Policies

Zirong Song, Zheng Lu, Haoran Liao, Wanqi Zhong, Yunhe Ni, Lijie Wang, Xingjie Fan, Zhisheng Chen, Yantang Qu, Meijia Chen, Tianyu Xin, Yiming Li, Xiuying Chen

arXiv 2609.06079首次发表:更新:

发表机构

Tsinghua University; MBZUAI; Nanyang Technological University; Zhejiang University; Rutgers University(清华大学; 穆罕默德·本·扎耶德人工智能大学; 南洋理工大学; 浙江大学; 罗格斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

LayerRoute提出动作条件混合层路由接口,自适应访问VLM层和动作表示,在LIBERO Long上提升7.2,仅增加少量参数。

AI 中文摘要

视觉-语言-动作(VLA)策略利用预训练的视觉-语言模型(VLM)来指导机器人控制中的动作生成。VLM提供跨层演化的分层视觉-语义表示,从局部视觉几何到抽象的、与语言对齐的语义;因此,不同的操作任务可能需要不同的层表示混合。同时,动作模块维护在动作计算过程中演化的中间表示,并可能为后续决策提供有用信息。然而,现有的VLA接口在表示访问方面提供的灵活性有限:VLM信息通过每个动作层的固定层分配来暴露,而中间动作状态仅通过残差流隐式传播,没有显式重用。我们引入LayerRoute,一种动作条件的表示路由接口,能够自适应访问VLM层和动作表示。层混合路由器动态形成缓存的VLM表示的混合,而动作状态重读则重用早期的动作表示。在多样化的模拟和真实世界基准测试中,LayerRoute一致地改进了StarVLA-$\pi$和$\pi_{0.5}$,在LIBERO Long上实现了高达7.2的提升,仅增加0.31% / 3.87%的参数。消融研究验证了动作条件层路由的益处,而路由分析揭示了跨动作层和任务设置的结构化分配模式。

英文摘要

Vision-Language-Action (VLA) policies leverage pretrained vision-language models (VLMs) to guide action generation for robot control. VLMs provide hierarchical visual-semantic representations that evolve across layers, from local visual geometry to abstract, language-aligned semantics; different manipulation tasks may therefore require different mixtures of layer representations. Meanwhile, the action module maintains intermediate representations that evolve throughout action computation and may provide useful information for subsequent decisions. However, existing VLA interfaces offer limited flexibility in representation access: VLM information is exposed through fixed layer assignments for each action layer, while intermediate action states are only propagated implicitly through residual streams without explicit reuse. We introduce LayerRoute, an action-conditioned representation routing interface that enables adaptive access to VLM layers and action representations. The Layer Mixture Router dynamically forms mixtures of cached VLM representations, while Action-State Reread reuses earlier action representations. Across diverse simulation and real-world benchmarks, LayerRoute consistently improves StarVLA-$π$ and $π_{0.5}$, achieving up to 7.2 gains on LIBERO Long with only 0.31% / 3.87% additional parameters. Ablation studies validate the benefit of action-conditioned layer routing, while routing analyses reveal structured allocation patterns across action layers and task settings.

Comments15 pages, 7 figures, 16 tables, including appendices

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑