arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.35469cs.ROcs.AIcs.LG

重新思考流匹配中的因果动作标记化与条件退火

Rethinking Causal Action Tokenization with Conditional Annealing in Flow Matching

  • IIIS, Tsinghua University(清华大学智能产业研究院)
  • Shanghai Qizhi Institute(上海期智研究院)

机构由 AI 辅助整理,请以论文原文为准。

Chenyu Zhang, Yuhang Cao, Daru Du, Yingxi Lu, Jing Shao, Ruoqu Chen, Jiajun Liu, Liu Cao, Yicheng Liu, Hang Zhao, Mengdi Xu

AI总结:

针对自回归视觉-语言-动作模型动作标记化与骨干网络语义不对齐的问题,提出因果动作标记器CATok,通过条件退火流匹配构建从粗到细的因果标记空间,配合MMDiT解码器实现高保真重建,在模拟和真实机器人任务中优于现有方法。

AI中文摘要:

自回归视觉-语言-动作(VLA)模型为机器人学习提供了一条可扩展的路径,然而现有的动作标记器将标记化视为一个压缩问题,产生的表示在语义上与自回归骨干网络不对齐。我们提出了CATok,一种因果动作标记器,它将标记化重新构建为一个因果结构的生成过程。CATok引入了一种条件退火机制,通过逐步退火流匹配过程来提取动作标记:每个标记都以所有先前的标记为条件,并在特定的噪声水平上编码残差重建信号,从而建立一个从粗到细的因果标记空间,其生成语义在结构上与自回归建模对齐。一个基于多模态扩散Transformer(MMDiT)的标记条件流匹配解码器,以混合扩散头架构的精度从这些离散标记中重建连续的动作块。这种离散瓶颈通过设计强制实现了知识隔离,清晰地将高层语义推理与低层运动执行分开,而无需显式的注意力掩码。在三个模拟基准和真实世界机器人操作任务上的广泛评估表明,CATok在重建保真度-压缩权衡和推理效率方面始终优于现有的标记化方法,同时提高了VLA任务成功率和训练效率,为纯自回归VLA系统建立了一个高性能、可扩展的基础。

英文摘要:

Autoregressive Vision-Language-Action (VLA) models offer a scalable path to robot learning, yet existing action tokenizers treat tokenization as a compression problem, producing representations that are semantically misaligned with the autoregressive backbone. We propose CATok, a causal action tokenizer that reframes tokenization as a causally structured generative process. CATok introduces a conditional annealing mechanism that extracts action tokens by progressively annealing a flow-matching process: each token is conditioned on all preceding tokens and encodes the residual reconstruction signal at a specific noise level, establishing a coarse-to-fine causal token space whose generative semantics are structurally aligned with autoregressive modeling. A token-conditioned flow-matching decoder built on Multimodal Diffusion Transformer (MMDiT) reconstructs continuous action chunks from these discrete tokens with the precision of hybrid diffusion-head architectures. This discrete bottleneck enforces knowledge insulation by design, cleanly separating high-level semantic reasoning from low-level motor execution without requiring explicit attention masking. Extensive evaluations across three simulation benchmarks and real-world robotic manipulation tasks demonstrate that CATok consistently surpasses existing tokenization methods in both reconstruction fidelity-compression tradeoff and inference efficiency, while improving VLA task success rate and training efficiency, establishing a high-performance, scalable foundation for purely autoregressive VLA systems.

↑