arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

StairVLA:面向视觉-语言-动作模型的阶段感知分层动作生成

StairVLA: Stage-Aware Hierarchical Action Generation for Vision-Language-Action Models

Shangyuan Yuan, Xinda Qi, Yujiang Pu, Wenliang Guo, Xiaobo Tan

arXiv 2610.07756首次发表:更新:

发表机构

Michigan State University; Ant Group(密歇根州立大学; 蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLA模型动作头统一去噪忽略阶段差异的问题,提出阶段感知分层生成框架StairVLA,以部分去噪动作为接口分层生成动作,在LIBERO上提升成功率并大幅降低推理延迟。

AI 中文摘要

视觉-语言-动作(VLA)模型日益依赖基于扩散或流匹配的动作头来生成连续的机器人动作。这些动作头通常以大致统一的方式处理去噪轨迹。然而,我们观察到,条件关注点会随去噪阶段自然转移:早期阶段结合语言指令和视觉观察以建立粗略的动作轨迹,而后期阶段则更强调当前的视觉观察以进行动作对齐。基于这一洞察,我们提出了StairVLA,一种阶段感知的分层动作生成框架,该框架利用部分去噪的动作作为粗略的长时程动作生成与局部细化之间的自然接口。一个高层VLA执行早期去噪以生成可复用的长时程部分去噪动作轨迹,而一个轻量级细化器则以更高频率运行,利用最新观察细化局部动作块。这种设计分摊了昂贵的高层VLA计算,同时保留了频繁的闭环校正。在LIBERO上,我们的GR00T风格实例将平均成功率从96.5%提升至97.8%,同时将每个动作块的摊销推理延迟从115.0毫秒降低至44.2毫秒。更广泛地,在两种VLA骨干网络、仿真基准和真实机器人任务中,StairVLA持续降低推理成本,同时保持强劲的任务性能。

英文摘要

Vision-language-action (VLA) models increasingly rely on diffusion- or flow-matching-based action heads to generate continuous robot actions. These action heads typically process the denoising trajectory in a largely uniform manner. However, we observe that the conditioning focus naturally shifts across denoising stages: early stages combine language instructions and visual observations to establish a coarse action trajectory, whereas later stages place greater emphasis on current visual observations for action alignment. Based on this insight, we introduce StairVLA, a stage-aware hierarchical action generation framework that uses partially denoised actions as a natural interface between coarse long-horizon action generation and local refinement. A high-level VLA performs early denoising to produce a reusable long-horizon partially denoised action trajectory, while a lightweight refiner operates at a higher frequency to refine local action chunks using the latest observations. This design amortizes expensive high-level VLA computation while preserving frequent closed-loop correction. On LIBERO, our GR00T-style instantiation improves average success from 96.5% to 97.8% while reducing amortized inference latency from 115.0 ms to 44.2 ms per action chunk. More broadly, across two VLA backbones, simulation benchmarks, and real-robot tasks, StairVLA consistently reduces inference cost while maintaining strong task performance.

CommentsProject page: https://dicomsky.github.io/projects/stairvla

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑