发表机构
University of California, Los Angeles; University of California, Santa Barbara; Amazon; Robotics and AI Institute(加利福尼亚大学洛杉矶分校; 加利福尼亚大学圣巴巴拉分校; 亚马逊公司; 机器人与人工智能研究院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出语义统一(SUN)程序及系统Kuafu,由大型视觉语言系统自动合成,在9项任务中实现82.03%宏观成功率,优于相关基准,可将控制摊销为鲁棒策略,实现符号规划与数据驱动执行的统一。
AI 中文摘要
在长程操作中,桥接基于模型的控制与学习到的策略存在一个潜在分歧:控制执行指定目标,学习将该行为摊销为反应式策略,但现有方案会丢弃任务语义,导致奖励需手动设计且行为与控制需求偏离。本文引入语义统一(Semantically UNified, SUN)程序,这是一种类型化可执行文件,其中几何和接触关系被定义一次,并编译为对齐的模型预测控制(Model Predictive Control, MPC)代价、满足谓词、强化学习(RL)奖励、转移守卫和诊断工具。本文的系统Kuafu由大型视觉语言系统驱动,可从语言和场景语义自动合成SUN程序,通过MPC筛选可行性,并在训练阶段条件策略时保留语义。在9项任务中,Kuafu实现了82.03%的宏观成功率,优于稀疏奖励基准(35.67%)和Stage-BC基准(24.75%);在8192规模下,其每小时成功轨迹时间为人类远程操作的10.57倍;每项任务使用500条轨迹时,Kuafu数据训练的DP3策略在模拟中达到46.0%的成功率,而替代方法为22.4%,在真实的Franka和Kinova机器人上则达到34.7%的成功率。这些结果表明,经模拟筛选的任务语义可有效将控制摊销为鲁棒策略,无需演示或手动密集奖励,实现符号规划与数据驱动执行的统一。
英文摘要
Model-based control can directly execute specified objectives, while learning can amortize such behaviors into reactive policies, making their combination a natural solution to multi-stage manipulation. We introduce Semantically UNified (SUN) Programs, typed executables that compile grounded relations into aligned optimal control objectives, satisfaction predicates, and learning rewards. Our harness, Kuafu, equips a foundation model as a task-level agent to orchestrate scene preparation, verification, residual RL, and data production. The agent uses program feedback to repair candidate programs and training diagnostics to calibrate relative reward weights, retaining accepted task semantics across tool calls. Across nine multi-stage manipulation tasks, Kuafu achieves 82.03% average success rate, significantly outperforming all learned baselines. Its learned controllers generate demonstrations at 10.57x the human-teleoperation rate, yielding data that improve visualpolicy success by 23.6 percentage points over the strongest baseline. The policies transfer zero-shot to physical Franka and Kinova robots, demonstrating sim-to-real generalization.