一类通过隐藏神经元的联想记忆中的相
Phases in a class of associative memories via hidden neurons
- CyberAgent AI Lab(CyberAgent AI实验室)
- RIKEN iTHEMS(RIKEN跨学科理论科学研究所)
- Graduate School of Artificial Intelligence and Science, Rikkyo University(立教大学人工智能与科学研究科)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文研究Krotov-Hopfield二部架构(类H),通过隐藏神经元作为有序参数,统一分析多项式与指数负载下的联想记忆相图,揭示可见与隐藏拉格朗日函数分别决定稳定性和存储规模。
AI中文摘要:
Hopfield网络中的联想记忆是无序多体系统中的吸引子动力学,高阶和指数扩展将其检索更新转变为softmax注意力机制。多项式和指数机制已通过不同方法分析,但缺乏一个共同的架构来探究什么决定了存储规模。本文研究了Krotov和Hopfield的二部架构,我们称之为类$H$,其模型由每层的拉格朗日函数确定,以隐藏神经元作为检索的有序参数。在多项式负载下,复制方法得到复制对称相图和闭式容量,且串扰矩对于Ising和球形可见神经元是共同的,因此它们的差异源于可见熵。对于softmax隐藏层,负载是指数级的,副本表示将热力学映射到随机能量模型计数,包含顺磁相、凝聚相和冻结相。加热通过注意力的量化重新分配使检索失稳,典型的Gaussian模式在每个负载下保持亚稳态。这些机制在串扰统计上不同,多项式负载下为中心极限,指数负载下为大偏差,类$H$将检索分为两个角色:可见拉格朗日函数决定稳定性,隐藏拉格朗日函数决定存储规模,这两个轴也可能指导新拉格朗日函数的设计。
英文摘要:
Associative memory in the Hopfield network is attractor dynamics in a disordered many-body system, and higher-order and exponential extensions turn its retrieval update into softmax attention. The polynomial and exponential regimes have been analyzed by different methods, with no common architecture in which to ask what fixes the storage scale. In this paper we study the bipartite architecture of Krotov and Hopfield, which we call the class $H$, whose model is fixed by a Lagrangian for each layer, taking the hidden neurons as the order parameter of retrieval. At polynomial load the replica method yields the replica-symmetric phase diagrams and closed-form capacities, and the crosstalk moment is common to Ising and spherical visible neurons, so their differences come from the visible entropy. With a softmax hidden layer the load is exponential, and a copy representation maps the thermodynamics onto random-energy-model counting, with paramagnetic, condensed, and frozen phases. Heating destabilizes retrieval by quantized reassignments of attention, and typical Gaussian patterns remain metastable at every load. The regimes differ in their crosstalk statistics, central-limit at polynomial load and large-deviation at exponential load, and the class $H$ splits retrieval into two roles, the visible Lagrangian fixing stability and the hidden one the storage scale, two axes that may also guide the design of new Lagrangians.