时间相关性如何塑造线性循环神经网络的记忆
How Temporal Correlations Shape Memory in Linear Recurrent Neural Networks
浏览论文内容
中文总结 AI 辅助
本研究针对线性循环神经网络,推导了相关输入下的学习动力学,揭示了时间相关性对记忆的影响,解释了相关数据使循环网络成为变化检测器的原因。
中文摘要 AI 辅助
线性循环神经网络(LRNN)是用于研究网络在训练过程中积累多少记忆的简单模型。对于不相关的输入,先前的研究发现,训练本身会使网络在保留过去信息与仅对当前信息做出反应之间达到平衡。而真实序列是相关的,我们针对相关输入精确求解了学习动力学。在该解中,保留过去信息会产生一种代价,相关性的全部影响都体现在这种代价上。当输入不相关时,这种代价会简化为先前研究中的代价;当输入呈正相关时,代价会增大。由此得出三个发现:(1)相关性不仅会改变学习的最终结果,还会重塑学习过程:记忆会积累、出现超调,且部分被消除,最终稳定的网络会保留更少的过去信息。(2)记忆会在一个阈值处关闭,该阈值由一个数值设定,即每个输入与紧邻其前一个输入的相似程度。序列长度或更长范围的相关性都不会改变这个阈值。只有当任务对前一个输入的需求超过当前输入通过与过去的相关性已提供的需求时,记忆才值得保留。(3)最优网络也会发生变化:零误差需要一个直通路径(feedthrough),该路径将当前输入直接传递到网络输出,无需记忆;当提供一个额外的隐藏维度时,训练会自动构建该路径。我们的研究将输入的一个属性转化为对网络是否学习记忆的预测,并解释了为何相关数据会将循环网络转变为变化检测器。
英文摘要
The linear recurrent neural network (LRNN) is a simple model for studying how much memory a network builds up as it trains. For uncorrelated inputs, earlier work found that training itself settles the network between keeping the past and reacting only to the present. Real sequences are correlated, and we solve the learning dynamics exactly for correlated inputs. In the solution, keeping the past carries a cost. The whole effect of correlation lands on that cost. This cost reduces to the earlier one when inputs are uncorrelated and grows once they are positively correlated. Three findings follow. (1) Correlation reshapes the course of learning, not only its end. Memory builds, overshoots, and is partly removed, and the settled network keeps less of the past. (2) Memory switches off at a threshold set by one number, how much each input resembles the one just before it. Neither sequence length nor longer-range correlation moves this threshold. Memory is worth keeping only when the task needs the previous input more than the current input already supplies it through correlation with the past. (3) The best network changes too. Zero error demands a feedthrough, a path that passes the current input straight to the network's output and remembers nothing, and training builds it unprompted when given one spare hidden dimension. Our work turns one property of the input into a prediction of whether a network learns memory and explains why correlated data turns recurrent networks into change detectors.
发表机构
- University of Buea(布埃亚大学)
- minoHealth AI Labs(minoHealth人工智能实验室)
机构由 AI 辅助整理,请以论文原文为准。