arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多层异联想神经网络中的指数容量

Exponential Capacity in Multilayer Hetero-Associative Neural Networks

Elena Agliari, Adriano Barra, Andrea Ladiana, Andrea Lepre

arXiv 2607.29554首次发表:更新:

AI 中文总结

该研究提出多层异联想指数神经网络,其存储模式数随层大小呈指数增长,在多领域数据上验证了该容量特性,揭示指数存储与强泛化为不同能力。

AI 中文摘要

指数Hopfield网络存储的模式数量随神经元数量呈指数增长,其经典形式为自联想网络,可将损坏的记忆副本补全为原始记忆。然而,人们希望这类网络执行的许多任务是异联想任务,即把线索映射到不同的目标。我们引入并分析了一个由L层N个二元神经元组成的指数神经网络,每层携带各自的数据集,其能量是各层Mattis重叠乘积的指数函数,因此仅当所有层检索到相同索引的模式时,能量才达到最小;存储的关联必须是线索的满射函数,我们证明了无法存储其他形式的关联。通过腔/信噪比分析(经噪声的大偏差评估在主导阶上精确),我们发现对齐的异联想态是零温动力学的不动点,可存储的模式数量P_c~e^{Nρ_L},该数量随层大小呈指数增长,其中显式速率ρ_L随Llog2增长;扩大吸引子盆地会降低速率,但不会破坏其指数特性。将该理论与结构化数据对比,我们发现指数容量和预测的吸引子盆地在相关的多对一模式下依然存在:该网络是近乎完美的内容可寻址记忆。相同的闭式表达式无需重新拟合即可描述合成流形、真实T细胞受体/表位三元组及自然语言意图数据,因此该机制具有领域普适性。对未见过线索的泛化性能虽显著高于随机水平,但仍低于记忆性能,且决定其超出随机水平程度的是编码几何而非数据领域。在该网络家族中,指数存储与强泛化是两种不同的能力。

英文摘要

Exponential Hopfield networks store a number of patterns that grows exponentially with the number of neurons, and in their classical formulation they are auto-associative: they complete a corrupted copy of a memory into the memory itself. Many of the tasks one wants such a network to perform are instead hetero-associative, mapping a cue to a different target. We introduce and analyse an exponential neural network of $L$ layers of $N$ binary neurons, each layer carrying its own dataset, whose energy is an exponential of the product of the per-layer Mattis overlaps, so that it is minimised precisely when every layer retrieves the pattern of the same index; the stored association must be a surjective function of the cue, and we show why nothing else can be stored at all. A cavity/signal-to-noise analysis, made exact at leading order by a large-deviation evaluation of the noise, shows that the aligned hetero-associative state is a fixed point of the zero-temperature dynamics up to a number of stored patterns $P_c\sim e^{Nρ_L}$, exponential in the layer size, with an explicit rate $ρ_L$ that grows like $L\log 2$; enlarging the basins of attraction lowers the rate but never destroys its exponential character. Comparing the theory with structured data we find that the exponential capacity and the predicted basins survive correlated, many-to-one patterns: the network is a near-perfect content-addressable memory. The same closed forms describe, without refitting, a synthetic manifold, real T-cell-receptor/epitope triples and natural-language intent data, so the mechanism is domain-universal. Generalisation to unseen cues, though significantly above chance, stays below memorisation, and it is the geometry of the encoding, rather than the data domain, that sets how far above chance it reaches. In this family, exponential storage and strong generalisation are distinct capabilities.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑