发表机构
Dovetail Research Group; ICC CONICET; Universidad de Buenos Aires(多尾研究集团; 阿根廷国家科学技术研究委员会国际合作中心; 布宜诺斯艾利斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文研究受控随机过程(含变换器)极小表示的扰动稳定性,引入近似同态概念,证明线性变换器在特定条件下的收敛性,为AI模型潜在表示的结构收敛假说提供理论支持。
AI 中文摘要
我们研究受控随机过程(特别是变换器)的极小表示在扰动下的稳定性,该问题的动机来自近期在神经网络的潜在表示中发现预测状态结构的实验。我们考虑标准变换器、线性变换器和预测变换器,引入近似同态的概念以捕捉它们之间的局部结构相似性,同时引入比较其诱导动力学(我们称之为接口)的度量,并证明近似同态的可组合性等性质。对于标准变换器,我们证明存在简单接口,使得动力学的不同实现之间不存在近似同态;相比之下,对于每个有限秩接口$\boldsymbol{\textit{I}}$,我们证明所有实现与$\boldsymbol{\textit{I}}$足够接近的接口的极小线性变换器,都存在到$\boldsymbol{\textit{I}}$的极小实现的近似同态,其误差与扰动大小呈线性关系。我们利用关于信念状态不可区分性的一些温和假设,在残差度量下为预测变换器证明了类似的稳定性结果。这些结果明确了规范变换器表示对扰动具有鲁棒性的条件,同时表明若没有额外的结构限制,这种收敛将无法成立。在这类抽象被嵌入现代AI模型的隐藏层的假设下,这为其潜在表示表现出结构收敛的假说提供了一些理论支持。
英文摘要
We study the stability of minimal representations of controlled stochastic processes (in particular, transducers) under perturbations. This question is motivated by recent experiments finding predictive-state structure in the latent representations of neural networks. We consider standard, linear and predictive transducers. We introduce notions of approximate homomorphism capturing local structural similarity between them, together with metrics comparing their induced dynamics (which we refer to as interfaces), and prove properties such as composability of the approximate homomorphisms. For standard transducers, we show that there exist simple interfaces for which there is no approximate homomorphism between the different implementations of the dynamics. In contrast, for every finite-rank interface $\mathcal I$, we prove that all minimal linear transducers implementing interfaces sufficiently close to $\mathcal I$ have an approximate homomorphism to the minimal implementation of $\mathcal I$, with error linear in the perturbation size. We prove an analogous stability result for predictive transducers under a residual metric using some mild hypothesis regarding the indistinguishability of the belief states. These results identify conditions under which canonical transducer representations are robust to perturbations, while showing that such convergence fails without additional structural restrictions. Under the assumption that these type of abstractions are embedded into the hidden layers of modern AI models, this gives some theoretical support to the hypothesis that their latent representations exhibit structural convergence.
Comments40 pages: 23 pages of main text and 17 pages of appendices; 6 figures