发表机构
University of Tennessee(田纳西大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究将LLM响应分类任务形式化为随机线性DS的二元假设检验,证明基于DS的分类误判率随序列长度指数衰减,还建立跨嵌入泛化的可迁移可区分性下界,解释了该方法的经验性能。
AI 中文摘要
近期研究表明,通过将令牌嵌入建模为黑盒动力系统(DS)的轨迹并比较两个DS的预测残差,可对大语言模型(LLM)的响应进行分类。尽管这种动力系统方法在经验上取得了成功,但对于其为何有效、如何随令牌序列规模扩展以及何时能跨嵌入模型迁移,仍缺乏理论理解。我们通过将分类任务形式化为两个随机线性DS之间的二元假设检验来解决这些问题。我们发现,即使两个DS的动态差异显著,它们的平稳边际分布之间的总变差距离也可以任意小,这为任何忽略令牌动态的分类器提供了基本精度下限。随后,我们证明基于DS的分类的误分类概率随序列长度L呈指数衰减,衰减由动态可区分性量δ²控制,该量捕获两个DS之间的谱距离。我们还通过引入嵌入模型之间的近似交织条件来表征跨嵌入泛化,并根据交织映射的最小奇异值建立可迁移可区分性的下界。这些结果共同解释了基于DS的分类的经验性能,并推动进一步研究用DS理论分析AI系统,这与用AI建模动力系统的更常见方法形成对比。
英文摘要
Recent work has shown that classifying large language models (LLMs)' responses can be distinguished by modeling token embeddings as trajectories of a black-box dynamical system (DS) and comparing prediction residuals of two DSs. Despite the empirical success of this dynamical approach, a theoretical understanding of why it works, how well it scales as a function of the token sequence, and when it transfers across embedding models remains lacking. We address these questions by formalizing the classification task as a binary hypothesis test between two stochastic linear DSs. We show that the total variation distance between the stationary marginal distributions of the two DSs can be arbitrarily small even when the dynamics differ substantially, which provides a fundamental accuracy floor for any classifier that ignores token dynamics. We then show that the misclassification probability of DS-based classification decays exponentially in the sequence length $L$, with the decay governed by a dynamical discriminability quantity $δ^2$ that captures the spectral distance between the two DSs. We also characterize cross-embedding generalization by introducing an approximate intertwining condition between embedding models and establishing a lower bound on the transferable discriminability in terms of the intertwining map's smallest singular value. Together, these results explain the empirical performance of DS-based classification and motivate further investigation into using DS theory to analyze AI systems, in contrast to the more common approach of using AI to model dynamical systems.