MAGMA-GEN:通过反事实重执行从模糊失败中验证恢复监督
MAGMA-GEN: Validated Recovery Supervision from Ambiguous Failures via Counterfactual Re-Execution
浏览论文内容
中文总结 AI 辅助
MAGMA-GEN通过反事实重执行验证恢复监督,将模糊失败转化为训练数据,提升长时程操作任务的成功率和恢复能力。
中文摘要 AI 辅助
执行长时程操作任务的分层机器人系统必须做出高层语义决策,以编排随机性的低层技能。在这种设置下,失败的轨迹回放是模糊的:一个较差的下游状态可能反映无效的高层决策、部分观测,或一个有效决策但其物理执行失败。传统的监督学习缺乏此类恢复状态的数据,而强化学习则面临稀疏奖励和非局部信用分配的问题。我们提出MAGMA-GEN,一种在策略数据生成流水线,将模糊的失败回放转换为经过验证的恢复监督。MAGMA-GEN首先使用一个特权教练来假设一个早期的决策级错误,并提出局部修正或恢复动作。由于这种诊断可能是错误的,候选动作仅当在匹配条件下从相同状态重新执行能改善下游进展时才被保留。这从智能体自身的失败分布中生成监督示例,无需逐步的人类演示。在交互式长时程操作任务上的评估表明,MAGMA-GEN在模拟和真实机器人执行中,面对不断变化的任务约束,相较于蒸馏和轨迹修复基线,提高了任务成功率和恢复能力。
英文摘要
Hierarchical robotic systems executing long-horizon manipulation tasks must make high-level semantic decisions that orchestrate stochastic low-level skills. In this setting, failed rollouts are ambiguous: a poor downstream state may reflect an invalid high-level decision, partial observation, or a valid decision whose physical execution failed. Traditional supervised learning lacks data for such recovery states, while reinforcement learning struggles with sparse rewards and non-local credit assignment. We propose MAGMA-GEN, an on-policy data-generation pipeline that converts ambiguous failed rollouts into validated recovery supervision. MAGMA-GEN first uses a privileged coach to hypothesize an early decision-level error and propose localized correction or recovery actions. Because this diagnosis is fallible, candidates are retained only if re-execution from the same state under matched conditions improves downstream progress. This produces supervised examples from the agent's own failure distribution without per-step human demonstrations. Evaluated on interactive long-horizon manipulation tasks, MAGMA-GEN improves task success and recovery capabilities, against distillation and trajectory-repair baselines under evolving task constraints in both simulation and real-robot execution.
发表机构
- LAAS-CNRS(法国国家科学研究中心分析与系统架构实验室)
机构由 AI 辅助整理,请以论文原文为准。