离散强制:将离散引导注入连续去噪以实现少步动作专家
Discrete Forcing: Infusing Discrete Guidance into Continuous Denoising for Few-Step Action Experts
- The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
- South China University of Technology(华南理工大学)
- University of Science and Technology of China(中国科学技术大学)
- Westlake University(西湖大学)
- Zhejiang University(浙江大学)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
提出离散强制框架,通过先预测离散动作令牌建立粗结构,再引导连续细化,实现少步高效动作生成,在多个基准和真实任务中性能与速度均优于连续专家。
中文摘要 AI 辅助
在视觉-语言-动作(VLA)模型中,高效的动作生成需要同时捕捉粗粒度的动作结构和细粒度的细节。离散动作令牌提供了紧凑的结构表示,但牺牲了精度,而连续动作令牌提供了高精度,但通常需要多个去噪步骤。我们提出了离散强制(Discrete Forcing),一种流匹配框架,通过显式的从粗到细的生成过程结合这两种表示。它首先预测离散动作令牌以建立粗粒度的动作结构,然后利用它们来引导连续动作的细化。离散和连续组件共享一个带有专门分支的通用扩散变压器骨干,保持了与常规单分支模型相当的参数数量,同时每个分支仅需一次前向传播。在多个基准上的广泛评估表明,与参数匹配的连续动作专家相比,性能有所提升,推理速度更快,并且随着模型容量的增加,性能持续提升。真实世界实验进一步展示了在高精度和动态操作任务上的改进。
英文摘要
Efficient action generation in vision-language-action (VLA) models requires capturing both coarse action structure and fine-grained details. Discrete action tokens provide compact structural representations but sacrifice precision, while continuous action tokens offer high precision but often require multiple denoising steps. We introduce Discrete Forcing, a flow-matching framework that combines these representations through an explicit coarse-to-fine generation process. It first predicts discrete action tokens to establish a coarse action structure, then uses them to guide continuous action refinement. The discrete and continuous components share a common diffusion transformer backbone with specialized branches, maintaining a parameter count comparable to a conventional single-branch model while requiring only one forward pass per branch. Extensive evaluations across multiple benchmarks demonstrate improved performance and faster inference over a parameter-matched continuous action expert, with consistent performance gains as model capacity increases. Real-world experiments further demonstrate improvements on high-precision and dynamic manipulation tasks.