AC-VLA:通过组合学习实现稳健的分布外动作执行
AC-VLA: Robust Out-of-Distribution Action Execution via Compositional Learning
浏览论文内容
中文总结 AI 辅助
研究VLA模型在分布外泛化的问题,提出AC-VLA框架,含组合学习模块和状态条件不对称掩码策略,无需修改架构可集成到VLA主干,在LIBERO和LIBERO-OOD基准测试中,该框架在组合OOD任务上有显著改进,分布内性能良好。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型在端到端机器人操作方面表现出色,但在分布外(OOD)泛化方面存在困难,当熟悉的子任务以未见的配置重新组合时。我们识别出两种相互强化的失败模式:轨迹过拟合,即模型过度拟合整体轨迹模式而非组合子技能语义;以及感知捷径,即动作令牌过度依赖手腕视图纹理而牺牲全局空间基础。为了解决这两个问题,我们引入了AC-VLA,一个即插即用的动作组合学习框架,包括两个与架构无关的组件:(i)一个组合学习模块,使用基于大语言模型的指令分解器和本体感觉轨迹对齐器生成密集的子任务监督,然后对完整演示和分解数据进行混合训练,赋予模型组合泛化能力;(ii)一种状态条件不对称掩码策略,在闭合抓取阶段抑制手腕视图输入,加强全局语义基础。所有组件无需架构修改,可直接集成到任何VLA主干中。在$\pi_{0.5}$上实例化并在LIBERO和LIBERO-OOD基准上评估,AC-VLA在组合OOD任务上实现了约28%的绝对改进,同时保持了近乎完美的分布内性能。
英文摘要
Vision-Language-Action (VLA) models excel at end-to-end robotic manipulation but struggle with out-of-distribution (OOD) generalization when familiar sub-tasks are recombined in unseen configurations. We identify two mutually reinforcing failure modes: \emph{trajectory overfitting}, where models overfit to holistic trajectory patterns rather than compositional sub-skill semantics; and \emph{perceptual shortcut}, where action tokens over-rely on wrist-view textures at the expense of global spatial grounding. To address both, we introduce \textbf{AC-VLA}, a plug-and-play Action Compositional learning framework comprising two architecture-agnostic components: \textbf{(i)} a compositional learning module that uses an LLM-driven instruction decomposer and a proprioceptive trajectory aligner to generate dense sub-task supervision, followed by mixed training on complete demonstrations and decomposed data to endow the model with compositional generalization; and \textbf{(ii)} a state-conditioned asymmetric masking strategy that suppresses wrist-view inputs during closed-gripper phases, enforcing global semantic grounding. All components are architectural modification-free and directly integrable into any VLA backbone. Instantiated on $π_{0.5}$ and evaluated on LIBERO and LIBERO-OOD benchmarks, AC-VLA achieves a ~28% absolute improvement on compositional OOD tasks while maintaining near-perfect in-distribution performance.
发表机构
- Shenzhen Technology University(深圳技术大学)
- Shenzhen University of Advanced Technology(深圳先进技术大学)
- Tongji University(同济大学)
- Great Bay University(大湾区大学)
机构由 AI 辅助整理,请以论文原文为准。