arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.18410cs.LG

用于高效视觉-语言-动作策略的角色条件子令牌路由

Role-Conditioned Sub-Token Routing for Efficient Vision-Language-Action Policies

  • Futurewei Technologies(未来智能技术公司)

机构由 AI 辅助整理,请以论文原文为准。

Wei Jiang, Wei Wang

AI总结:

该研究针对VLA模型推理成本高的问题,提出RoleSub方法,通过角色条件子令牌路由压缩视觉和语言表示,在匹配视觉-KV预算下多数场景优于仅令牌控制,结合压缩后总KV仅为原始的9.2-11.3%且控制性能强。

AI中文摘要:

视觉-语言-动作(VLA)模型处理长多模态令牌序列,导致推理在内存和计算上都成本高昂。现有的效率方法主要是减少视觉令牌,但激进的令牌剪枝会变得脆弱,因为移除一个令牌会丢弃其整个表示。子令牌压缩提供了一种补充替代方案,它在保留更多令牌的同时降低它们的值宽度。然而,直接将子令牌压缩应用于VLA策略效果较差,因为对感知、语言理解和控制重要的信息在多模态表示中的分布不同。我们引入角色条件子令牌路由(RoleSub),它学习如何压缩保留令牌的值表示。在视觉令牌减少后,RoleSub在正交空间中将每个保留的值表示划分为组,并使用轻量级路由确定应保留哪些组。路由决策取决于令牌表示、学习到的潜在角色表示和语言上下文。相同的机制也可应用于语言值,允许压缩视觉和语言表示而不移除额外的令牌。我们在OpenVLA-OFT-7B上跨四个LIBERO套件评估RoleSub。在匹配的视觉-KV预算下,RoleSub在36个设置中的33个中优于仅训练令牌的控制,在激进压缩下获得最大收益。结合视觉和语言压缩将总KV降低到原始的9.2-11.3%,同时在大多数任务上保留强大的控制性能。这些结果表明,减少保留令牌内的表示为激进的VLA压缩提供了对令牌剪枝的有效补充。

英文摘要:

Vision-Language-Action (VLA) models process long multimodal token sequences, making inference expensive in both memory and computation. Existing efficiency methods mainly reduce visual tokens, but aggressive token pruning becomes fragile because removing a token discards its entire representation. Sub-token compression provides a complementary alternative by retaining more tokens while reducing their value width. However, directly applying sub-token compression to VLA policies is less effective because information important for perception, language understanding, and control is distributed differently across the multimodal representation. We introduce Role-Conditioned Sub-Token Routing (RoleSub), which learns how to compress the value representations of retained tokens. After visual token reduction, RoleSub partitions each retained value representation into groups in an orthogonal space and uses a lightweight router to determine which groups should be preserved. The routing decision is conditioned on the token representation, a learned latent role representation, and language context. The same mechanism can also be applied to language values, allowing visual and language representations to be compressed without removing additional tokens. We evaluate RoleSub on OpenVLA-OFT-7B across the four LIBERO suites. At matched visual-KV budgets, RoleSub outperforms a trained token-only control in 33 of 36 settings, with the largest gains under aggressive compression. Combining visual and language compression reduces total KV to 9.2--11.3% of the original while retaining strong control performance on most tasks. These results show that reducing the representation within retained tokens provides an effective complement to token pruning for aggressive VLA compression.

补充信息

↑