arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09454cs.RO

RobotAPO:面向机器人操作视频生成的对抗性物理偏好优化

RobotAPO: Adversarial Physics Preference Optimization for Robotic Manipulation Video Generation

Kerui Li, Zhe Jing, Chenyi Huang, Xiaofeng Wang, Zheng Zhu, Haoming Cui, Huaibo Huang

首次发表
浏览论文内容

中文总结 AI 辅助

提出RobotAPO框架,利用对抗性偏好优化在流匹配去噪空间中纠正机器人操作视频的局部物理违规,提升物理一致性并显著改善真实机器人任务成功率。

中文摘要 AI 辅助

机器人操作视频越来越多地被用作具身代理的视觉规划,但仅优化视觉合理性往往无法捕捉真实世界交互中脆弱的物理流形。即使在交互边界处出现微小的违反物理规律的错误,如相互穿透或物体过早运动,也可能完全破坏下游执行所需的推断时序和姿态。由于标准的监督微调缺乏直接惩罚这些局部失效的压力,我们引入了AgiBot-PhysPref。这个经过严格筛选的10,000样本偏好数据集隔离了条件匹配的物理违规,将生成器自身的失败分布转化为物理一致性的基础信号。在此基础上,我们提出了RobotAPO,一种在连续流匹配去噪空间中运行的对抗性物理偏好优化框架。为了防止策略仅仅记忆静态的精选失败,RobotAPO采用了一个轻量级的对抗性反事实提议器,该提议器在去噪空间中学习一个依赖于条件的、偏向物理失败的方向。这鼓励模型探索并更好地尊重物理交互边界,同时保持纯粹的提示和参考推理接口,无需外部结构条件。综合评估表明,显式纠正这些局部物理违规可改善从生成视频中进行的下游机器人执行。在保留的AgiBot条件下,RobotAPO在物理一致性方面比最强的受控内部基线高出6.8%的硬分数和10.0%的软分数。至关重要的是,在真实机器人重放中,它将这些物理一致性增益转化为相对于最强受控内部基线任务成功率37.4%的相对提升。

英文摘要

Robotic manipulation videos are increasingly used as visual plans for embodied agents, but optimizing purely for visual plausibility often fails to capture the fragile physical manifold of real-world interactions. Even minor physics-violating errors at the interaction boundary, such as interpenetration or premature object motion, can completely invalidate the inferred timing and pose needed for downstream execution. Because standard supervised fine-tuning lacks the direct pressure to penalize these localized failures, we introduce AgiBot-PhysPref. This rigorously curated 10,000-sample preference dataset isolates condition-matched physics violations, turning the generator's own failure distribution into a foundational signal for physical consistency. Building upon this, we propose RobotAPO, an adversarial physics preference optimization framework operating in the continuous flow-matching denoising space. To prevent the policy from merely memorizing static curated failures, RobotAPO employs a lightweight adversarial counterfactual proposer that learns a condition-dependent, physical-failure-biased direction in denoising space. This encourages the model to explore and better respect the physical interaction boundary, all while maintaining a pure prompt-and-reference inference interface without requiring external structural conditioning. Comprehensive evaluations demonstrate that explicitly correcting these localized physics violations improves downstream robot execution from generated videos. On held-out AgiBot conditions, RobotAPO outperforms the strongest controlled internal baseline in physical consistency by 6.8% hard score and 10.0% soft score. Crucially, in real-robot replay, it translates these physical-consistency gains into a 37.4% relative improvement in task success over the strongest controlled internal baseline.

发表机构

  • Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)
  • GigaAI(极佳科技)
  • Beijing Institute of Technology(北京理工大学)

机构由 AI 辅助整理,请以论文原文为准。

↑