发表机构
Pengcheng Laboratory(鹏城实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对VLA实时控制延迟问题,提出紧迫性感知去噪框架,按动作紧迫性分配计算,实现高达1.89倍延迟加速且成功率相当。
AI 中文摘要
扩散和流匹配视觉-语言-动作(VLA)策略通过迭代去噪生成动作块,导致严重的推理延迟,极大限制了实时机器人控制。现有加速方法将动作块视为一个整体计算单元,忽略了后退时域控制的一个关键物理现实:动作是联合生成但顺序消费的,导致执行紧迫性固有地异构。我们利用这种不对称性,提出了紧迫性感知去噪(UAD),一种新颖的推理时框架,根据每个动作物理上需要的时间分配去噪计算。UAD在更少的去噪步骤后释放时间紧迫的动作,同时将尾部动作的持续背景细化与物理执行重叠。然而,异构去噪引入了两个关键挑战:紧迫动作的早期释放错误和尾部动作的轨迹不一致性。UAD通过两个核心机制优雅地解决了这两个问题:轨迹协调,它重建统一的内部状态演化以恢复联合去噪一致性,无需额外模型评估;以及幽灵动作校正,它利用未执行的幽灵延续动态补偿剩余可执行动作的早期释放错误。跨多个VLA架构、模拟基准和真实世界操作任务的广泛评估表明,UAD在平均动作可用性延迟上实现了高达1.89倍的加速,同时保持了与具有最佳去噪预算的普通推理相当的成功率,提供了比最先进的VLA加速基线更优的成功-延迟权衡。
英文摘要
Diffusion and flow-matching Vision-Language-Action (VLA) policies generate action chunks through iterative denoising, incurring substantial inference latency that severely limits real-time robotic control. Existing acceleration methods treat an action chunk as a monolithic computational unit, ignoring a crucial physical reality of receding-horizon control: actions are generated jointly but consumed sequentially, resulting in inherently heterogeneous execution urgencies. We exploit this asymmetry to introduce Urgency-Aware Denoising (UAD), a novel inference-time framework that allocates denoising computation according to when each action is physically needed. UAD releases time-critical urgent actions after fewer denoising steps while overlapping the continued background refinement of tail actions with physical execution. However, heterogeneous denoising introduces two key challenges: early-release errors in urgent actions and trajectory inconsistency in tail actions. UAD elegantly resolves both through two core mechanisms: Trajectory Reconciliation, which reconstructs unified internal state evolution to restore joint denoising coherence without additional model evaluations, and Ghost Action Correction, which leverages non-executed ghost continuations to dynamically compensate for early-release errors across remaining executable actions. Extensive evaluations across multiple VLA architectures, simulation benchmarks, and real-world manipulation tasks demonstrate that UAD achieves up to a 1.89x speedup in average action availability latency while maintaining comparable success rates to vanilla inference with optimal denoising budget, offering a more favorable success-latency trade-off than state-of-the-art VLA acceleration baselines.
Comments19 pages