发表机构
University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对长程机器人技能衔接处的观测空间偏移问题,构建检测-恢复-重启系统,在BOSS-44基准上将全链成功率从7.6%提升至26.5%,验证了场景错位是故障主因。
AI 中文摘要
长程机器人操作通常通过串联独立训练的技能构建。尽管每个技能单独使用时可靠,但串联后性能会急剧下降:每个下游技能必须从前序技能留下的状态开始,而非其训练分布。我们研究这种故障模式——观测空间偏移(Observation-Space Shift, OSS),并探究技能衔接处故障的原因。利用特权模拟器重置,我们发现主要偏移来自场景状态的错位(例如前序技能留下的打开的抽屉或次要物体),而非机器人关节配置或下游技能操作的物体。为验证该诊断,我们构建了一个完全学习的检测-恢复-重启系统:任务进度监视器检测停滞,学习策略恢复错位的场景组件,衔接处鲁棒微调使技能重启。该系统在所有测试替代方法失败的衔接处实现恢复,我们将此视为诊断的证据而非通用方法。在BOSS-44基准上,该系统将全链成功率从7.6%提升至26.5%,较基础策略提升3.5倍,达到特权恢复神谕的51%,而最佳K重采样、扩散策略(Diffusion Policy)和世界模型基线均无法从评估的衔接状态中恢复。在运行微调π₀.₅策略的真实Franka机械臂上,同一监视器受外部相机可观测性限制,但闭环仍可恢复部分原本的终端故障,这推动了腕部和夹爪传感的发展。这些结果表明,部分长程组合故障通过在重启策略前恢复场景,比从非支持状态重试更易解决。
英文摘要
Long-horizon robotic manipulation is often built by chaining independently trained skills. Although each skill can be reliable in isolation, performance degrades sharply when skills are chained: each downstream skill must start from the state its predecessor leaves behind rather than from its training distribution. We study this failure mode, Observation-Space Shift (OSS), and ask what causes these skill-seam failures. Using privileged simulator resets, we find that the dominant shift comes from displaced scene state (e.g., an open drawer or secondary objects left behind by earlier skills), not from the robot's joint configuration or the object the downstream skill manipulates. To test this diagnosis, we build a fully learned detect-restore-resume system: a task-progress monitor detects the stall, a learned policy restores the displaced scene components, and seam-robust fine-tuning lets the skill resume. It recovers the seam where every tested alternative fails, which we treat as evidence for the diagnosis rather than as a general-purpose method. On the BOSS-44 benchmark, the system improves full-chain success from 7.6% to 26.5%, a 3.5x improvement over the base policy and 51% of a privileged restoration oracle, whereas best-of-K resampling, a Diffusion Policy, and world-model baselines fail to recover from the evaluated seam states. On a real Franka arm running a fine-tuned $π_{0.5}$ policy, the same monitor is limited by exterior-camera observability, yet closing the loop still recovers some otherwise-terminal failures, motivating wrist and gripper sensing. These results suggest that some long-horizon composition failures are better addressed by restoring the scene before resuming the policy than by retrying from an off-support state.
Comments8 pages, 4 figures, 7 tables. Submitted to ICRA 2027