发表机构
The University of Sydney(悉尼大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出WA-SpecDec框架,在VLA预填充阶段注入世界模型的物理场景感知,在不改变放宽接受规则的前提下,提升了任务成功率、实现1.5倍匹配成功率加速并降低近接触失败率。
AI 中文摘要
视觉-语言-动作(VLA)策略以自回归方式生成机器人控制指令,导致闭环延迟主要由目标模型的多次前向传播决定。投机解码通过并行验证草稿动作令牌块降低该成本,近期VLA方法进一步放宽了令牌级接受规则,因为动作令牌空间的微小差异通常对应相似的连续控制量。然而,这种放宽仍与场景无关:固定的令牌距离容忍度在不同状态下对相同的动作令牌偏差视为同等安全,尽管在自由空间中无害的偏差在接触区域附近可能导致碰撞或抓取失败。本文提出WA-SpecDec,一种世界感知的投机解码框架,该框架在VLA预填充阶段注入世界模型衍生的物理场景感知,为草稿提议和目标验证生成共享的世界感知预填充状态,且不改变放宽的接受规则。在三种最先进的放宽接受方案中,WA-SpecDec在更宽松的接受规则下保持了更高的任务成功率,并支持更长的接受前缀。在可比成功率的操作点下,WA-SpecDec相较于单独的VLA投机解码实现了1.5倍的匹配成功率加速,且相较于对应的投机基线平均降低了18.6%的近接触失败(NCF)率。
英文摘要
Vision-language-action (VLA) policies generate robot controls autoregressively, making closed-loop latency dominated by repeated target-model forward passes. Speculative decoding reduces this cost by verifying blocks of draft action tokens in parallel, and recent VLA methods further relax token-level acceptance because small differences in action-token space often map to similar continuous controls. However, this relaxation remains scene-agnostic. A fixed token-distance tolerance treats the same action-token deviation as equally safe across states, although deviations that are harmless in free space can cause collisions or grasp failures near contact. We propose WA-SpecDec, a world-aware speculative decoding framework that injects world-model-derived physical scene awareness during the VLA prefill stage, producing shared world-aware prefill states for draft proposal and target verification without changing the relaxed acceptance rule. Across three state-of-the-art relaxed acceptance schemes, WA-SpecDec preserves higher task success under looser relaxation and enables longer accepted prefixes. At comparable-success operating points, WA-SpecDec achieves a 1.5x matched-success speedup over VLA speculative decoding alone and reduces near-contact failure (NCF) by 18.6% on average relative to the corresponding speculative baselines.
CommentsPreprint