发表机构
Stanford University; Columbia University(斯坦福大学; 哥伦比亚大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究运动条件机器人协同设计问题,提出基于RoboTokens训练的Transformer Transformer统一模型,能跨实体空间和用例,通过动力学自引导优化设计,实验显示其可零样本优化,制造的优化设计大幅降低跟踪误差。
AI 中文摘要
机器人操作性能中一个常被忽视的因素是机器人自身的实体体现。受此问题启发,我们研究运动条件机器人协同设计,目标是生成完整的机器人设计,在优化用户定义奖励的同时跟踪(来自人类示范的)目标末端执行器轨迹。我们引入Transformer Transformer,这是一种基于RoboTokens训练的扩散变换器,RoboTokens是机器人实体、状态和动作的统一令牌化表示。相同架构可用于不同实体空间(如轮式双臂、四足动物、类人机器人)和用例(实体生成、跨实体控制器)。Transformer Transformer是一个动力学模型,其与奖励无关的状态和动作预测可转换为特定奖励的价值预测,通过我们称为动力学自引导的过程来引导实体扩散向高价值机器人设计。跨多个设计空间的实验显示了对未见奖励和轨迹的零样本优化,比进化基线提高了性能和运行时效率。最后,我们制造了一个优化的ALOHA设计,与原始设计相比,跟踪误差降低了70%以上。
英文摘要
An often overlooked factor of robot manipulation performance is the embodiment of the robot itself. Motivated by this problem, we study motion-conditioned robot co-design, where the goal is to generate complete robot designs that track target end-effector trajectories (from human demonstrations) while optimizing user-defined rewards. We introduce Transformer Transformer, a diffusion transformer trained on RoboTokens, a unified tokenization of robot embodiments, states, and actions. The same architecture can be used across embodiment spaces (e.g., wheeled bimanual, quadrupeds, humanoids) and use cases (embodiment generation, cross embodiment controller). Rather than overfitting to one reward function, Transformer Transformer is a dynamics model, whose reward-agnostic state and action predictions can be converted into reward-specific value predictions. These value predictions are used to steer embodiment diffusion towards high value robot designs, through a procedure we call Dynamics Self-Guidance. Experiments across multiple design spaces show zero-shot optimization of unseen rewards and trajectories, improving performance and runtime over the evolutionary baseline. Finally, we fabricated an optimized ALOHA design, which reduced tracking error by over 70% compared to the original design.
Comments26 pages, 12 figures, 20 tables. Project page: https://transformer-transformer.github.io