LexiconVLA:为未见任务学习可复用的原子动作码本
LexiconVLA: Learning Reusable Atomic Action Codebooks for Unseen Tasks
- Sun Yat-sen University(中山大学)
- X-Era AI Lab(X-Era AI实验室)
- Pengcheng Laboratory(鹏城实验室)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
提出LexiconVLA,通过全局与细节码本捕获可复用交互结构与执行细节,结合轨迹重建和视觉对齐学习原子动作词典,实现未见任务跨任务复用,在RLBench和真实机器人上显著提升成功率。
AI中文摘要:
视觉-语言-动作(VLA)模型在未见任务中难以复用重复出现的交互。我们的诊断性研究表明,可靠的任务完成并不意味着一系列原子动作在不同任务情境下被一致地执行。我们提出LexiconVLA,一个可检索的原子动作词典,用于跨任务复用。全局码本和细节码本分别捕获共享的交互结构和细粒度的执行变化,从而保留可复用模式和执行细节。视觉-原子动作对齐将视觉状态变化中的轨迹重建与动作码中的视觉结果预测相结合,将词典建立在运动及其效果之上。我们利用包含69个任务中57,803个片段的AtomAction数据集,通过轨迹重建和视觉对齐来学习这些码本。一个规划器和场景适配器将新目标转化为由码条件化的子任务,供共享策略使用,无需特定技能的专家或部署时的参数更新。在26个RLBench任务上,跨五个策略主干,LexiconVLA在18个已见任务上基本保持性能,同时在8个从策略训练中保留的任务上提高了成功率。使用BridgeVLA,未见任务的成功率从16.67%提升至34.17%(+17.50个百分点),总体成功率达到了71.08%,是报告结果的方法中最高的。真实机器人实验展示了逐步执行和失败恢复能力。
英文摘要:
Vision-language-action (VLA) models struggle to reuse recurring interactions in unseen tasks. Our diagnostic study reveals that reliable task completion does not imply consistent execution of constituent atomic actions across task contexts. We present LexiconVLA, a retrievable atomic-action lexicon for cross-task reuse. Global and detail codebooks capture shared interaction structure and fine-grained execution variation, respectively, preserving both reusable patterns and execution details. Visual-Atomic Action Alignment couples trajectory reconstruction from visual state changes with visual outcome prediction from action codes, grounding the lexicon in motion and its effects. We learn these codebooks with trajectory reconstruction and visual alignment on our AtomAction Dataset of 57,803 segments from 69 tasks. A planner and scene-aware adapter translate new goals into code-conditioned subtasks for a shared policy, without skill-specific experts or deployment-time parameter updates. Across five policy backbones on 26 RLBench tasks, LexiconVLA largely maintains performance on 18 seen tasks while improving success on 8 tasks held out from policy training. With BridgeVLA, unseen-task success rises from 16.67% to 34.17% (+17.50 percentage points), and overall success reaches 71.08%, the highest among methods with reported results. Real-robot experiments demonstrate stepwise execution and failure recovery.