网络的训练历史何时比其当前状态更能预测其未来学习?来自响应探针和预测屏幕的证据
When does a network's training history predict its future learning better than its current state? Evidence from a response probe and a forecasting screen
- Technische Universität Ilmenau(伊尔梅瑙工业大学)
- German Centre for Integrative Biodiversity Research (iDiv) Halle–Jena–Leipzig(德国综合生物多样性研究中心(iDiv)哈勒-耶拿-莱比锡)
- Friedrich Schiller University(弗里德里希·席勒大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究通过响应探针和预测屏幕实验,探究网络训练历史何时比当前状态更能预测未来学习,发现历史仅在当前状态尚未提供信息时才有预测价值。
AI中文摘要:
当前行为相似的网络在继续训练时仍可能以不同方式学习。关于可塑性丧失和关键期的工作表明,通往状态的路径塑造了后续发展;但并未表明该路径是否携带了状态本身测量所遗漏的信息。我们探究网络的训练历史何时比其当前状态更能预测其未来学习。在一项主要研究中,小型多层感知器在三种历史机制(42种历史)下训练,并通过一个短探针在四个检查点测量未来学习:该探针是一个在新任务上训练100次更新的网络副本。在读取预测结果之前,协议检查了探针。它对隐藏单元的功能保持性重新缩放作出单调响应,重复测量一致(在可靠性最低的类别中,三次重复均值的组内相关系数为0.940,[0.903, 0.997]),并且单元的重新初始化在其后立即可见,但在100至200次更新后不可见。一个最多四维的历史状态并未优于当前状态的校准模型(增益-21.4%,90%区间[-91.9, 8.1];预先要求:10%)。一项对1,560次合成回归运行的伴随屏幕对更远目标(运行的最终误差)提出了相同问题。在那里,历史模型在12个(最多240个时期中的)时期后比当前验证误差预测得更好(紧凑状态30.3%,[15.8, 39.4],一种情境比较),并且在48个时期后与之无法区分。在两项研究中,历史仅在当前状态尚未对目标提供信息时才有信息价值;这一解读是在结果之后形成的。
英文摘要:
Networks that behave alike now can still learn differently when training continues. Work on loss of plasticity and critical periods shows that the path to a state shapes what follows; it does not show whether the path carries information that a measurement of the state itself misses. We ask when the training history of a network predicts its future learning better than its current state. In a main study, small multilayer perceptrons were trained under three history regimes (42 histories), and future learning was measured at four checkpoints by a short probe: a copy of the network trained for 100 updates on a new task. Before the prediction result was read, the protocol checked the probe. It responded monotonically to a function-preserving rescaling of hidden units, repeated measurements agreed (intraclass correlation 0.940, [0.903, 0.997], in the least reliable class, mean of three repeats), and a re-initialisation of units was visible directly after it but not 100 to 200 updates later. A history state of at most four dimensions did not improve on a calibrated model of the current state (gain -21.4%, 90% interval [-91.9, 8.1]; required in advance: 10%). A companion screen on 1,560 synthetic regression runs asked the same question for a target further away, the final error of the run. There, history models forecast better than the current validation error after 12 of up to 240 epochs (compact state 30.3%, [15.8, 39.4], a contextual comparison) and were not distinguishable from it after 48. In both studies the history was informative only while the current state was not yet informative about the target; this reading was formed after the results.