arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16153cs.RO

用于精准单步动作生成的统一条件-动作建模

Unified Condition-Action Modeling for Accurate One-Step Action Generation

  • NTU(南洋理工大学)
  • HKU(香港大学)

机构由 AI 辅助整理,请以论文原文为准。

Xinyu Zhou, Zikun Cai, Kuangji Zuo, Gen Li, Boyu Ma, Yanshuo Lu, Yutong Song, Mingqi Yuan, Jiayu Chen, Jianfei Yang

AI总结:

该研究提出UCA-Flow统一条件-动作建模框架,通过共享令牌空间与改进双路监督方案,提升机器人单步动作生成的成功率与推理速度。

AI中文摘要:

机器人操纵需要兼具精准性与高效性的策略,因为机器人控制必须在严格延迟约束下对变化的观测做出响应。近期的扩散策略和流策略颇具潜力,但它们常将条件视为辅助信号,而非与动作轨迹共同演化。我们发现,这一局限可通过一种简单却有效的统一条件-动作建模设计得到有效缓解,该设计将条件与动作表示在共享令牌空间中,使紧凑模型能实现高性能,同时提升推理速度与精准度。因此,我们提出UCA-Flow,这是一种用于精准单步动作生成的统一条件-动作建模框架。我们的方法将观测条件、时间步条件、区间条件及动作令牌统一为单序列,并通过统一条件-动作Transformer处理,以实现联合条件-动作表示学习。结果是,条件表示会根据当前生成阶段动态重构,突出与动作优化最相关的信息。此外,我们引入了一种针对u和v的改进双路监督方案,以更强地优化统一条件-动作建模。UCA-Flow相比最强基线将平均成功率提升了9.3个百分点,同时相比DP3和Simple DP3分别实现了45.6倍和33.4倍的加速,且分别比单步FlowPolicy和MP1快4.3倍和2.3倍。

英文摘要:

Robot manipulation requires policies that are both accurate and efficient, as robot control must respond to changing observations under tight latency constraints. Recent diffusion and flow policies are promising, but they often treat conditions as auxiliary signals rather than jointly evolving them with action trajectories. We find that this limitation can be effectively mitigated by a \textbf{simple yet effective unified condition-action modeling design} that represents conditions and actions in a shared token space, allowing a compact model to achieve high performance while improving both inference speed and accuracy. Therefore, we propose UCA-Flow, a unified condition-action modeling framework for accurate one-step action generation. Our method unifies observation conditions, timestep conditions, interval conditions, and action tokens into a single sequence, and processes them with a Unified Condition-Action Transformer for joint condition-action representation learning. As a result, condition representations are dynamically reconstructed according to the current generation stage, highlighting information most relevant for action refinement. Furthermore, we introduce an improved dual-pass supervision scheme over $u$ and $v$ for stronger optimization of unified condition-action modeling. UCA-Flow improves the average success rate by 9.3 percentage points over the strongest baseline, while achieving $45.6\times$ and $33.4\times$ speedups over DP3 and Simple DP3, and remaining $4.3\times$ and $2.3\times$ faster than one-step FlowPolicy and MP1, respectively.Project page: https://uca-policy.github.io/UCA.github.io/.

↑