arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39526cs.RO

离散强制:将离散引导注入连续去噪以实现少步动作专家

Discrete Forcing: Infusing Discrete Guidance into Continuous Denoising for Few-Step Action Experts

  • The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
  • South China University of Technology(华南理工大学)
  • University of Science and Technology of China(中国科学技术大学)
  • Westlake University(西湖大学)
  • Zhejiang University(浙江大学)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Jingbo Wang, Wenxuan Song, Wenhao Yu, Han Zhao, Xi Wang, Jiayi Chen, Donglin Wang, Yan Wang, Haoang Li

中文总结 AI 辅助

提出离散强制框架,通过先预测离散动作令牌建立粗结构,再引导连续细化,实现少步高效动作生成,在多个基准和真实任务中性能与速度均优于连续专家。

中文摘要 AI 辅助

在视觉-语言-动作(VLA)模型中,高效的动作生成需要同时捕捉粗粒度的动作结构和细粒度的细节。离散动作令牌提供了紧凑的结构表示,但牺牲了精度,而连续动作令牌提供了高精度,但通常需要多个去噪步骤。我们提出了离散强制(Discrete Forcing),一种流匹配框架,通过显式的从粗到细的生成过程结合这两种表示。它首先预测离散动作令牌以建立粗粒度的动作结构,然后利用它们来引导连续动作的细化。离散和连续组件共享一个带有专门分支的通用扩散变压器骨干,保持了与常规单分支模型相当的参数数量,同时每个分支仅需一次前向传播。在多个基准上的广泛评估表明,与参数匹配的连续动作专家相比,性能有所提升,推理速度更快,并且随着模型容量的增加,性能持续提升。真实世界实验进一步展示了在高精度和动态操作任务上的改进。

英文摘要

Efficient action generation in vision-language-action (VLA) models requires capturing both coarse action structure and fine-grained details. Discrete action tokens provide compact structural representations but sacrifice precision, while continuous action tokens offer high precision but often require multiple denoising steps. We introduce Discrete Forcing, a flow-matching framework that combines these representations through an explicit coarse-to-fine generation process. It first predicts discrete action tokens to establish a coarse action structure, then uses them to guide continuous action refinement. The discrete and continuous components share a common diffusion transformer backbone with specialized branches, maintaining a parameter count comparable to a conventional single-branch model while requiring only one forward pass per branch. Extensive evaluations across multiple benchmarks demonstrate improved performance and faster inference over a parameter-matched continuous action expert, with consistent performance gains as model capacity increases. Real-world experiments further demonstrate improvements on high-precision and dynamic manipulation tasks.

↑