arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SparkVLA:用于长程操作的停止感知分层视觉-语言-动作(VLA)模型,结合自适应动作分块

SparkVLA: Stop-Aware Hierarchical VLA with Adaptive Action Chunking for Long-Horizon Manipulation

Xunyao Lei, Renjun Wu, Tianlin Huo, Xuesong Li

arXiv 2608.16172首次发表:更新:

AI 中文总结

针对分层VLA系统中停止与动作分块决策孤立评估的问题,提出SparkVLA,通过统一排序解决依赖,在RoboCerebra上取得显著性能提升,真实机器人实验验证了其有效性。

AI 中文摘要

在分层视觉-语言-动作(VLA)系统的每个重观测点,必须做出两个接口决策:何时终止当前子任务,以及执行提议的动作分块的长度。这些决策相互依赖——最优停止点取决于执行器计划执行的内容,而最优执行长度取决于子任务边界所在位置——但现有架构却孤立地评估这两个决策,这种不对称性是任何单个模块都无法克服的。我们提出SparkVLA,这是一种停止感知分层VLA模型,通过将两个决策表述为单一排序来解决这种相互依赖关系:停止决策与统一候选集中的每个动作前缀长度竞争,系统选择得分最高的选项,消除了阈值调整需求,仅需离线序数偏好。锚点条件上下文编码模块缓存感知历史的子任务锚点编码、起始状态记忆和目标语义,引导视觉标记向任务相关区域修剪;停止感知动作前缀选择头在分块边界通过全自注意力对所有候选进行评分,以保证效率。在RoboCerebra数据集上,SparkVLA达到47.12%的成功率,超过官方分层基准30.57%,超过最强可复现方法26.83%。多步任务的真实机器人实验进一步在物理硬件上验证了这些性能提升。

英文摘要

At every re-observation point in a hierarchical Vision-Language-Action (VLA) system, two interface decisions must be made: when to terminate the current subtask and how far to execute the proposed action chunk. These decisions are mutually dependent---the optimal stopping point depends on what the executor plans to do, while the optimal execution length depends on where the subtask boundary lies---yet existing architectures evaluate them in isolation, an asymmetry neither module can overcome alone. We present SparkVLA, a stop-aware hierarchical VLA that resolves this mutual dependency by formulating both decisions as a single ranking: Stop competes against every action-prefix length in a unified candidate set, and the system selects the highest-scoring option, eliminating threshold tuning and requiring only offline ordinal preferences. An Anchor-Conditioned Context Encoding module caches a history-aware subtask anchor encoding onset-state memory and goal semantics, guiding visual-token pruning toward task-relevant regions; a Stop-Aware Action-Prefix Selection head scores all candidates via full self bnattention at chunk boundaries for efficiency. On RoboCerebra, SparkVLA achieves 47.12% success rate, surpassing the official hierarchical baseline by 30.57% and the strongest reproducible method by 26.83% Real-robot experiments on multi-step tasks further validate these gains on physical hardware.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑