用于物理人机交互中实时在线运动学习的自由能门控可塑性
Free-Energy-Gated Plasticity for Real-Time Online Motor Learning in Physical Human-Robot Interaction
- Okinawa Institute of Science and Technology Graduate University(冲绳科学技术大学院大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
该研究针对物理人机交互的实时在线运动学习问题,扩展PV-RNN并提出FEGP方法,实验证实其能提升运动模式覆盖与保留能力,关键在于可塑性的时间分配而非平均幅度。
AI中文摘要:
完全在线具身学习要求在持续交互过程中,既通过突触适应获取新行为,又保留先前学习的动力学特性。我们扩展了受预测编码启发的变分循环神经网络(PV-RNN),以持续调整其突触权重,并提出自由能门控可塑性(FEGP),该方法根据变分自由能调节有效学习率。在实时物理人机交互中,随机初始化的网络无需离线预训练、重放或任务边界信号,就获得了三种循环运动模式,且所有三种模式均在自主滚出中出现。针对每个流包含10个随机教学流和5种网络初始化的受控实验显示,FEGP大幅提升了 repertoire 覆盖范围,以及在先前获取的模式离开近期观测窗口后的保留能力。与门控时间平均有效率匹配的恒定学习率,或时间组织被打乱的相同增益值重放,均无法复现这些改进。这些结果表明,相对于模型-环境失配的可塑性时间分配,而非其平均幅度或分布,对持续在线学习过程中维持先前获取的行为至关重要。
英文摘要:
Fully online embodied learning requires synaptic adaptation to acquire new behaviors while preserving previously learned dynamics during ongoing interaction. We extend the Predictive-Coding-inspired Variational Recurrent Neural Network (PV-RNN) to continuously adapt its synaptic weights and propose Free-Energy-Gated Plasticity (FEGP), which regulates the effective learning rate according to variational free energy. In real-time physical human-robot interaction, a randomly initialized network acquired three cyclic motor patterns without offline pretraining, replay, or task-boundary signals, with all three patterns emerging in autonomous rollouts. Controlled experiments over ten randomized teaching streams and five network initializations per stream showed that FEGP substantially improved repertoire coverage and retention of previously acquired patterns after they left the recent observation window. Neither a constant learning rate matched to the gate's time-averaged effective rate nor replay of the same gain values with disrupted temporal organization reproduced these improvements. These results indicate that the temporal allocation of plasticity relative to model-environment mismatch, rather than simply its average magnitude or distribution, is critical for maintaining previously acquired behaviors during continued online learning.