面向机器人操作的轨迹级连续动作表示
Trajectory-Level Continuous Action Representation for Robotic Manipulation
浏览论文内容
中文总结 AI 辅助
该研究提出CAT轨迹级连续动作表示框架,通过频率感知位置编码等技术解决现有视觉运动系统的动作表示冗余问题,在LIBERO等数据集上显著提升机器人操作任务成功率。
中文摘要 AI 辅助
我们提出了CAT,这是一种面向机器人操作的轨迹级连续动作表示框架。现有视觉运动系统常将动作表示与控制频率绑定,或依赖固定的时间参数化方式,这会在高采样率下产生表示冗余,并限制对关键运动的建模。CAT则将固定实时区间内的动作轨迹编码为一组连续的隐式token。为确保不同控制频率下的时间一致性,我们进一步引入了频率感知位置编码,以建立共享的时间坐标系。轨迹级正则化进一步稳定了隐式表示。该方法避免了表示量随时间步密度增长,且无需依赖预定义的时间参数化方式。在LIBERO、MimicGen及真实世界长时程操作任务上开展的大量系统级评估表明,在匹配的训练设置下,基于CAT的策略始终优于具有竞争力的基于VQ和连续视觉运动的基线方法。在不同模型骨干和控制频率下,CAT均能稳定提升成功率。这些结果凸显了轨迹级连续动作建模在适配不同控制频率的可扩展机器人操作中的优势。
英文摘要
We propose CAT, a trajectory-level continuous action representation framework for robotic manipulation. Existing visuomotor systems often entangle action representation with control frequency or rely on fixed temporal parameterizations. This leads to representational redundancy at high sampling rates and limits the modeling of critical motion. CAT instead encodes action trajectories within a fixed real-time interval into a set of continuous latent tokens. To ensure temporal consistency across varying control frequencies, we further incorporate a frequency-aware positional encoding that establishs a shared temporal coordinate system. Trajectory-level regularization further stabilizes the latent representation. This approach prevents representation growth with timestep density and avoids reliance on predefined temporal parameterizations. Extensive system-level evaluations on LIBERO, MimicGen, and real-world long-horizon manipulation tasks demonstrate that CAT-based policies consistently outperform both competitive VQ-based and continuous visuomotor baselines under matched training settings. Across various model backbones and control frequencies, CAT consistently improves success rates. These results highlight the advantages of trajectory-level continuous action modeling for scalable robotic manipulation across varying control rates.
发表机构
- Fudan University(复旦大学)
- TeleAI, China Telecom(中国电信TeleAI)
机构由 AI 辅助整理,请以论文原文为准。