StellaVLA:用于通用视觉-语言-动作模型的上下文结构化演示
StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models
浏览论文内容
中文总结 AI 辅助
StellaVLA是一种测试时基于结构化演示适应的VLA框架,通过双训练设计提升泛化性,在VLA-Arena等基准上表现优于现有模型,可助力VLA模型适应OOD任务。
中文摘要 AI 辅助
视觉-语言-动作(VLA)模型可遵循指令并操作物体,但当场景、视角或物体与训练数据不同时,其性能往往会在分布外(OOD)情况下崩溃。为适应每种新场景,通常需要收集更多数据并进行微调。我们提出StellaVLA,这是一个在测试时通过基于单个检索到的演示进行条件调整来实现适应的框架。其核心思想是超越模仿专家的操作,转而传递操作的原因:一条自动化离线流水线将每个原始轨迹转换为结构化演示,例如任务计划、子目标描述和口头化的3D运动,且无需人工标注成本。该结构化演示作为上下文指导提供,使策略能够对任务进行推理而非模仿像素轨迹,这也使其可在不同实体(真实机器人、人手或XR演示)间迁移。并行双训练设计通过联合动作与语言目标在训练期间内化这种推理,而推理时仅使用动作专家,保持了实时高频控制且无额外延迟。在2026年8月1日的VLA-Arena排行榜上,StellaVLA以0.63的总体评分位列第一,而强基线模型(π₀.₅和LingBot-VLA)的评分分别为0.44和0.22;在LIBERO数据集上,其平均成功率达98.8%,在LIBERO-Plus数据集上成功率达85.1%。我们的真实机器人基准测试表明,StellaVLA可将人类/机器人演示以及人到机器人(XR)演示用作上下文结构化演示,帮助VLA模型适应分布外任务。
英文摘要
Vision-Language-Action (VLA) models can follow instructions and manipulate objects, but their performance often collapses out of distribution (OOD), when the scene, viewpoint, or object differs from training. Adapting to each new situation typically requires collecting more data and fine-tuning. We present StellaVLA, a framework that instead adapts at test time by conditioning on a single retrieved demonstration. The key idea is to move beyond imitating what an expert did and instead convey why: an automated offline pipeline converts each raw trajectory into a structured demonstration, e.g., a task plan, sub-goal descriptions, and verbalized 3D motion, at zero human-annotation cost. Provided as in-context guidance, this structured demonstration lets the policy reason about the task rather than mimic a pixel trajectory, which also makes it transferable across embodiments (real-robot, human-hand, or XR demonstrations). A parallel dual-training design internalizes this reasoning during training through a joint action-and-language objective, while inference uses the action expert alone, preserving real-time, high-frequency control with no added latency. On the VLA-Arena leaderboard(Aug 1, 2026), StellaVLA ranks first with an overall score of 0.63, versus 0.44 and 0.22 for the strong prior models ($π_{0.5}$ and LingBot-VLA), and it further leads on LIBERO with 98.8% average success rate and LIBERO-Plus with 85.1% success rate. Our real-robot benchmark demonstrates that StellaVLA can use both human/robot demos and human-to-robot (XR) demos as in-context structured demonstration to help VLA model adapt to OOD tasks.