arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自修复循环集成用于分布漂移的实时恢复

Self-Repairing Recurrent Ensembles for Real-Time Recovery from Distribution Shift

Julian Lemmel, Pedro D. Wendel Garcia, Taisuke Kobayashi, Radu Grosu

arXiv 2610.03249首次发表:更新:

发表机构

Vienna University of Technology; Medical University of Vienna; National Institute of Informatics(维也纳工业大学; 维也纳医科大学; 国立信息学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出一种由循环网络集成构成的控制器,通过随机掩码观测和卡尔曼融合生成自监督标签,在线微调各成员以应对传感器漂移等分布变化,实现无需监督的实时恢复。

AI 中文摘要

将预训练控制器部署到实际环境中,会使其面临训练数据中不存在的情况。传感器漂移、传感器完全故障以及累积的测量噪声都会引发分布漂移,从而可能使原本胜任的策略崩溃;通常这种情况发生在没有专家能够提供纠正标签的时刻。我们提出了一种方法,使策略能够在线且无需监督地从此类漂移中恢复。我们的控制器是一个循环网络集成,每个网络观察观测向量的随机掩码子集,其高斯输出通过顺序卡尔曼融合进行组合,使得置信度高的成员主导共识动作。在部署时,我们将此共识视为自监督标签,并针对该标签对每个成员进行微调,将每个成员的贡献按其卡尔曼增益平方的补数进行缩放。梯度使用RFLO计算,这是一种高效且生物上合理的实时循环学习近似方法,因此每次环境步骤后都会进行参数更新,策略在漂移发生时即作出反应。在一系列模拟连续控制任务中,我们的方法在传感器漂移后恢复接近原始性能,而看到完整观测的集成则无法恢复。同一框架还涵盖了完全在线的交互式模仿学习:当专家在场时,共识标签被专家动作替换,相同的更新规则在遥操作期间优化策略。

英文摘要

Deploying a pretrained controller exposes it to conditions that are absent from its training data. Sensor drift, outright sensor failure and accumulating measurement noise all induce a distribution shift that can collapse an otherwise competent policy; typically at a point in time where no expert is available to supply corrective labels. We present a method that lets a policy recover from such shifts online and without supervision. Our controller is an ensemble of recurrent networks, each of which observes a randomly masked subset of the observation vector, and whose Gaussian outputs are combined through sequential Kalman fusion so that confident members dominate the consensus action. At deployment, we treat this consensus as a self-supervised label and fine-tune each member towards it, scaling each member's contribution proportional to the complement of its squared Kalman gain. Gradients are computed using RFLO, an efficient and biologically plausible approximation of Real-Time Recurrent Learning, so that a parameter update follows every environment step and the policy reacts to a shift as it unfolds. On a range of simulated continuous control tasks, our approach recovers close to the original performance after a sensor shift, while ensembles that see the full observation are unable to recover. The same framework subsumes fully online interactive imitation learning: when an expert is present, the consensus label is replaced by the expert action and the identical update rule refines the policy during teleoperation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑