arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

线性循环记忆足以提炼机器人空气曲棍球的World-Model策略

Linear Recurrent Memory Suffices to Distil a World-Model Policy for Robot Air Hockey

F. Olivia Fan, Oliver Obst

arXiv 2609.39151首次发表:更新:

发表机构

UNSW Sydney(新南威尔士大学悉尼分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过提炼DreamerV3教师模型至64维线性循环策略,证明在机器人空气曲棍球防守中,线性记忆足以应对跟踪丢失,非线性循环并非必需。

AI 中文摘要

记忆依赖的控制是否需要非线性循环动力学?我们研究了在跟踪丢失情况下模拟空气曲棍球防守任务。DreamerV3教师模型在跟踪丢失下的表现优于无记忆策略,而重置教师的循环状态会急剧降低性能,这表明该任务需要记忆。我们将该教师模型提炼为具有64维状态的紧凑循环策略,结合对角线性循环和可选的秩-$k$非线性创新,同时保留非线性观测编码器和动作头。在五个匹配的随机种子下,纯线性循环模型($k=0$)在整个测试的跟踪丢失范围内与GRU基线和教师模型均匹配。增加非线性创新秩未带来可测量的收益。该结果是在新的测试分割上获得的,该分割将在所有模型和分析冻结后才开放。线性模型比GRU需要更少的循环参数和计算量,但性能相当。这些结果表明,对于这种记忆依赖的控制任务,围绕简单线性记忆机制的非线性表示学习可能就足够了,而非线性循环动力学并非必需。这些结论仅限于所考虑的模拟任务、教师模型、状态维度和黑屏时长,以及观测编码器和动作头保持非线性的策略。

英文摘要

Does memory-dependent control need nonlinear recurrent dynamics? We study simulated air-hockey defence under temporary loss of puck tracking. A DreamerV3 teacher outperforms a memoryless policy under tracking loss, while resetting the teacher's recurrent state sharply reduces performance, which demonstrates that the task requires memory. We distil this teacher into compact recurrent policies with a 64 dimensional state, with a combination of a diagonal linear recurrence and an optional rank-$k$ nonlinear innovation while retaining nonlinear observation encoders and action heads. Across five matched seeds, the purely linear recurrent model ($k=0$) matches both the GRU baseline and the teacher throughout the tested range of tracking loss. Increasing nonlinear innovation rank providing no measured benefits. This result is obtained on a fresh test split, which will be only opened after all models and analyses are frozen. The linear model requires fewer recurrent parameters and less computation than GRU, but performs comparably. These results suggest that, for this memory dependent control task, nonlinear representation learning around a simple linear memory mechanism can be sufficient, and that nonlinear recurrent dynamics are not necessarily required. These conclusions are limited to the simulated task, teacher, state dimension, and blackout horizon considered here, and to policies whose observation encoder and action head remain nonlinear.

Comments10 pages, 4 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑