arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.23452cs.ROcs.AIcs.LG

面向弹性空间机器人的无奖励持续适应

Reward-Free Continual Adaptation for Resilient Space Robots

发表机构卢森堡大学
查看机构详情
  • University of Luxembourg(卢森堡大学)

机构由 AI 辅助整理,请以论文原文为准。

Andrej Orsula, Miguel Olivares-Mendez, Carol Martinez

首次发表
浏览论文内容

中文总结 AI 辅助

针对空间机器人硬件退化导致传统控制策略失效的问题,提出无奖励持续学习框架,利用隐态世界模型使智能体无需奖励即可适应环境,在三类模拟任务中验证了方法有效性。

中文摘要 AI 辅助

空间机器人在极端环境中运行,硬件退化会严重损害传统控制策略。持续强化学习为在线适应提供了有前景的机制,但其在部署过程中固有地需要奖励信号。然而,由于缺乏外部跟踪系统及环境整体复杂性,在太空中精确计算奖励通常不可行。为解决不可观测奖励的挑战,我们引入一种利用隐态世界模型的无奖励持续学习框架。通过在多样模拟中预训练基于模型的智能体,世界模型在其隐空间内学习到奖励结构的鲁棒预测器。在部署到存在严重硬件退化的环境后,我们冻结观测编码器和奖励预测器,仅通过无监督回滚更新世界模型的转移动力学。通过完全基于该更新世界模型生成的想象轨迹训练策略,智能体无需接收新奖励即可适应改变的动力学。我们在遭受严重形态故障的模拟行星遍历、轨道导航和精密装配任务中验证了所提方法。

英文摘要

Space robots operate in extreme environments where hardware degradation can critically compromise traditional control strategies. While continual reinforcement learning offers a promising mechanism for online adaptation, it inherently requires access to a reward signal during deployment. However, precise reward computation in space is often infeasible due to the lack of external tracking systems and the overall complexity of the environment. To address the challenge of unobservable rewards, we introduce a reward-free continual learning framework that leverages latent-state world models. By pre-training a model-based agent across diverse simulations, the world model learns a robust predictor of the reward structure within its latent space. Upon deployment to an environment with severe hardware degradation, we freeze the observation encoder and reward predictor to update only the transition dynamics of the world model through unsupervised rollouts. By training the policy entirely on imagined trajectories generated by this updated world model, the agent adapts to altered dynamics without receiving new rewards. We demonstrate our approach across simulated planetary traversal, orbital navigation, and precision assembly tasks subjected to severe morphological failures.

补充信息

↑