AI 中文总结
针对具身AI的数据瓶颈,提出Pegasus低资源框架,通过结构化知识传递将人类操作视频转化为机器人可学习数据,经多基准与机器人评估验证其跨具身翻译及数据生成的有效性。
AI 中文摘要
具身人工智能的关键瓶颈并非模型架构,而是数据。尽管网络上存在数十亿个人类操作视频,但由于人类形态与机器人硬件之间的具身差距,机器人无法直接从中学习。我们引入Pegasus,一种低资源框架,通过结构化知识传递弥合这一差距,将人类演示转化为机器人可学习的数据。Pegasus不依赖原始视频提示,而是构建基于图的中间表示:从人类视频中提取的任务图,通过 affordance(可供性)图和约束图转换为机器人规划图,用于机器人条件下的视频生成。分层可供性潜在空间对物体状态、可供性和任务之间的关系进行建模,实现了超出物体身份的泛化。闭环物理验证器还利用运动学可行性、碰撞约束和关节极限过滤无效生成内容。我们在一系列以自我为中心的操作基准(包括GTEA Gaze+和EPIC-KITCHENS-100)以及多种机器人具身上评估Pegasus,评估指标包括任务正确性、可执行性、状态一致性和可学习性。结果表明,Pegasus可实现可靠的跨具身翻译,并证明机器人数据生成可从硬件收集问题重新定义为可扩展的低资源知识传递问题。
英文摘要
The key bottleneck in embodied AI is not model architecture but data. Although billions of human manipulation videos exist online, robots cannot directly learn from them due to the embodiment gap between human morphology and robot hardware. We introduce Pegasus, a low-resource framework that bridges this gap by translating human demonstrations into robot-learnable data through structured knowledge transfer. Instead of relying on raw video prompts, Pegasus constructs a graph-based intermediate representation: a Task Graph extracted from human videos is transformed through Affordance and Constraint Graphs into a Robot Planning Graph for robot-conditioned video generation. A hierarchical affordance latent space models the relationship between object states, affordances, and tasks, enabling generalization beyond object identities. A closed-loop physics verifier further filters invalid generations using kinematic feasibility, collision constraints, and joint limits. We evaluate Pegasus across a range of egocentric manipulation benchmarks, including GTEA Gaze+ and EPIC-KITCHENS-100, and diverse robot embodiments, assessing Task Correctness, Executability, State Consistency, and Learnability. Results demonstrate reliable cross-embodiment translation and show that robot data generation can be reframed from a hardware collection problem into a scalable, low-resource knowledge transfer problem.