发表机构
Universitat Pompeu Fabra; Institute for Cross-Disciplinary Physics and Complex Systems (IFISC), CSIC–UIB; Reed College(庞培法布拉大学; 跨学科物理与复杂系统研究所(IFISC),西班牙国家研究委员会-巴伊亚群岛大学; 里德学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文通过扰动连接组储层计算架构,发现性能与超参数鲁棒性存在权衡,该权衡与归一化前原始谱半径相关,且秀丽隐杆线虫连接组处于低方差区间。
AI 中文摘要
储层计算为研究递归网络架构如何塑造计算提供了一个受控环境:输入信号通过一个固定的非线性动力系统被投影到高维状态空间中,仅训练读出层。然而,储层性能可能依赖于超参数;本文探讨哪些递归网络特征支持对这些参数变化的鲁棒性。我们使用记忆容量(MC)、截断单延迟信息处理容量(IPC)和核秩(KR)来表征计算性能。跨输入历史的泛化能力通过泛化秩(GR)来衡量,而超参数鲁棒性则通过目标谱半径、输入缩放、泄漏率和神经元偏置扫描中各指标的变异系数(CV)来量化。为了考察鲁棒性的架构决定因素,我们构造了改变连接拓扑、兴奋/抑制符号结构、权重幅度和权重位置同时保留互补属性的扰动。在这些实验中,秀丽隐杆线虫连接组始终处于相对低方差区间。核心结果是性能-鲁棒性权衡:具有更高任务无关性能的架构变体也往往表现出更大的超参数敏感性和更差的公共尾部泛化。在兴奋/抑制边平衡扫描和洗牌对照中,这种权衡与归一化前的原始谱半径密切相关。由于每个扰动矩阵都被重新缩放到相同的目标半径,原始谱半径较低的矩阵其递归权重获得更大的全局放大。因此,观察到的架构间差异表征了谱半径归一化下结构变异和架构特定全局重新缩放的联合效应。
英文摘要
Reservoir computing provides a controlled setting for studying how recurrent network architectureshapes computation: input signals are projected into a high-dimensional state space by a fixed nonlinear dynamical system, and only the readout is trained. However, reservoir performance can be dependent on hyperparameters; this paper asks which recurrent network features support robustness to those parameter changes. We characterize computational performance using memory capacity (MC), truncated single-delay information-processing capacity (IPC), and kernel rank (KR). Generalization across input histories is measured using generalization rank (GR), while hyperparameter robustness is quantified using the coefficient of variation (CV) of each metric across sweeps of target spectral radius, input scaling, leak rate, and neuron bias. To examine the architectural determinants of robustness, we construct perturbations that alter connectivity topology, excitatory/inhibitory sign structure, weight magnitudes, and weight placement while preserving complementary properties. Across these experiments, the C. elegans connectome consistently occupies a relatively low-variance regime. The central result is a performance-robustness tradeoff: architecture variants with higher task-agnostic performance also tend to exhibit greater hyperparameter sensitivity and poorer common-tail generalization. Across the E/I edge balance sweeps and shuffle controls, this tradeoff is closely associated with the raw spectral radius before normalization. Because every perturbed matrix is rescaled to the same target radius, matrices with lower raw spectral radius receive greater global amplification of their recurrent weights. The observed differences among architectures therefore characterize the joint effects of structural variation and architecture-specific global rescaling under spectral-radius normalization.
Comments20 pages, 12 figures, including supplementary material. Submitted to Neural Networks