发表机构
Nanyang Technological University; Guangzhou Cloudbutterfly Technology Co., Ltd.(南洋理工大学; 广州云蝶科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
FIERCE框架利用通用策略初始化,通过统一的进展-失败评估器提供反馈,在有限交互下精炼紧凑专家策略,无需模拟器或密集奖励,在模拟和真实接触任务中验证了高效性与部署成本。
AI 中文摘要
通用机器人策略提供了有用的初始化,但通过有限物理交互来精炼紧凑专家策略需要信息丰富的学习反馈。我们提出FIERCE,一个以通用策略初始化的强化学习框架,其核心是一个统一的、任务自适应的进展-失败评估器。该架构在观察进展头与动作条件潜在预测器之间共享观察-语言表示,后者的过去和当前预测馈送给因果序列头以进行任务失败估计。来自进展和偏好标签、同步命令与观察以及终端结果的联合监督训练评估器;目标任务回放支持适应和校准。固定的评估器快照提供进展塑形和失败风险惩罚,以及独立验证的终端奖励,而评估器和策略更新在收集新经验时交替进行。精炼过程既不需要持续的通用策略动作查询,也不需要专用的目标任务模拟器或手动标注的密集奖励。部署时仅保留紧凑专家策略。评估在模拟和两个接触丰富的真实任务中分离了反馈质量、策略学习效率和部署成本。代码、模型权重和数据恢复工具已在https网址发布。
英文摘要
Generalist robot policies offer useful initialization, but refining compact specialists through limited physical interaction requires informative learning feedback. We present FIERCE, a generalist-initialized reinforcement learning framework centered on a unified, task-adaptive progress-failure evaluator. Its architecture shares an observation-language representation between an observed-progress head and an action-conditioned latent predictor whose past and current predictions feed a causal sequence head for task-failure estimation. Joint supervision from progress and preference labels, synchronized commands and observations, and terminal outcomes trains the evaluator; target-task rollouts support adaptation and calibration. Fixed evaluator snapshots provide progress shaping and failure-risk penalties alongside independently verified terminal rewards, while evaluator and policy updates alternate as new experience is collected. Refinement requires neither continued generalist action queries nor a dedicated target-task simulator or manually annotated dense rewards. Only the compact specialist is retained at deployment. The evaluation separates feedback quality, policy-learning efficiency, and deployment cost across simulation and two contact-rich real tasks. Code, model weights, and data-restoration tools are released at https://github.com/ar-mine/FIERCE.
Comments8 pages, 5 figures