使用复合算子线性化LLM语义变换
Using Composition Operators to Linearize LLM Semantic Transformations
- Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出利用复合算子(推广的Koopman算子)将LLM语义变换近似为矩形无限维算子,揭示其等距结构,并构造有限维近似以通过奇异值谱比较任务和模型。
AI中文摘要:
机器学习学习函数:从提示到响应,从图像到标题。这些函数在数学上是什么仍然难以言说。我们提出了一种方法,利用动力系统中属于Koopman主义范畴的技术来近似这类变换。我们引入了复合算子的使用,它推广了Koopman算子,并且关键的是,可以在不同空间之间进行映射,这促使了将LLM变换视为矩形无限维算子的观点。这种形式化揭示了有用的结构:在提示和响应分布的自然假设下,LLM算子是一个等距算子,学习表示之间的错位表现为其有限截断的谱污染。然后,我们概述了一种构造LLM算子有限维近似的方法,并演示了如何使用奇异值谱来比较任务和模型。
英文摘要:
Machine learning learns functions: prompt to response, image to caption. What these functions are mathematically remains hard to say. We present a method to approximate these kinds of transformations using techniques from dynamical systems that fall under the umbrella of Koopmanism. We introduce the use of composition operators, which generalize the Koopman operator and, crucially, can map between distinct spaces, motivating the perspective that LLM transformations are rectangular infinite-dimensional operators. This formalism reveals useful structure: under natural assumptions on the prompt and response distributions, the LLM operator is an isometry, and misalignment between learned representations manifests as spectral pollution of its finite sections. We then outline a method of constructing finite-dimensional approximations of an LLM operator, and demonstrate how the singular value spectrum can be used to compare tasks and models.