arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.03872cs.RO

EvoHIL:用于鲁棒交互式强化学习的自演化奖励与流匹配策略优化

EvoHIL: Self-Evolving Reward and Flow-Matched Policy Optimization for Robust Human-in-the-Loop Reinforcement Learning

Shuoqin Zhang, Tongtong Cheng, Xiru Gao, Jinzhuo Peng, Bin Zheng, Jiahao Tu, Ke Wang, Jia Pan, Zhe Hu, Kai Liu

首次发表
浏览论文内容

中文总结 AI 辅助

EvoHIL是适配奖励模型、动作生成器和视觉域的统一交互式学习框架,在六类机械臂操纵任务中提升了成功率、运动平滑度等性能。

中文摘要 AI 辅助

交互式强化学习(HIL-RL)使机器人能从有限的现实世界交互中学习接触丰富的操纵任务,但部署时暴露出三个耦合的局限:静态视觉奖励模型在场景变化下失效;独立采样的动作导致时间上不一致的运动;基于视觉的策略对外观变化仍敏感。我们提出EvoHIL,这是一个在分阶段交互式学习过程中适配奖励模型、动作生成器和视觉域的统一框架。首先,自演化奖励(SER)从人类确认的正样本和临时弱负样本中适配成功分类器。其次,动作流稳定(AFS)通过流匹配生成时间一致的动作块,将策略更新建立在已执行的动作前缀和示范行为之上。第三,感知保留的离线微调重放相关交互数据,同时将AFS的演员-评论家锚定到先前行为,无需额外机器人交互即可适配视觉域。在Franka FR3和SO-101机械臂的六个操纵任务上,于受控光照变化条件下,EvoHIL相较于交互式和模仿学习方法,提升了任务成功率、与人类确认标签的一致性、运动平滑度和完成时间。

英文摘要

Human-in-the-loop reinforcement learning (HIL-RL) enables robots to learn contact-rich manipulation from limited real-world interaction, but deployment exposes three coupled limitations: static visual reward models fail under scene changes; independently sampled actions cause temporally inconsistent motion; and vision-based policies remain sensitive to appearance shifts. We present EvoHIL, a unified framework that adapts the reward model, action generator, and visual do main within a staged human-in-the-loop learning process. First, self-evolving reward (SER) adapts the success classifier from human-confirmed positives and provisional weak negatives. Second, Action Flow Stabilization (AFS) generates temporally coherent action chunks through flow matching, grounding policy updates in executed action prefixes and demonstrated behavior. Third, retention-aware offline fine-tuning replays relit interaction data while anchoring the AFS actor-critic to prior behavior, adapting the visual domain without additional robot interaction. Across six manipulation tasks on Franka FR3 and SO-101 arms under a controlled lighting shift, EvoHIL improves task success, agreement with human-confirmation labels, motion smoothness, and completion time relative to human-in-the-loop and imitation baselines.Project page: https://anonymous4366.github.io/EvoHIL/

↑