发表机构
Innopolis University; Federal University of Ceará; Systems Research Institute of Polish Academy of Science; Warsaw University of Technology(因诺波利斯大学; 塞阿拉联邦大学; 波兰科学院系统研究所; 华沙理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本综述以两种互补视角组织LLM的张量方法,提供统一符号与理论基础,分析Transformer组件的张量化策略,引入ρ_gap指标,为张量化语言模型提供结构化入门路径。
AI 中文摘要
大型语言模型(LLM)由词元表示、权重、适配更新、缓存和激活等结构化高维对象构建而成,其多重线性结构未被以矩阵为中心的传统视角充分利用。张量分解与张量网络为该结构提供了一种原则性代数语言,但现有文献常将其视为孤立的压缩机制。本综述通过两种互补视角组织LLM的张量方法:一是涵盖分词、嵌入、预训练、适配、压缩、推理与可解释性的七阶段生命周期分类法,二是涵盖嵌入、注意力与前馈网络的组件视角。我们提供统一符号与理论基础,分析单个Transformer组件的张量化策略,比较各生命周期阶段的方法,同时明确评估协议与模型规模的差异。我们进一步将张量方法与邻近的效率技术及概率张量网络关联,最后综合开放挑战并引入ρ_gap指标——用于衡量理论内存减少与实测系统级加速之间的压缩实现差距。通过将张量化视为通用结构原则,本综述为张量化语言模型提供了结构化入门路径,并阐明参数节省何时可合理转化为内存效率、计算效率或可解释性。本论文的GitHub页面可通过指定URL访问。
英文摘要
Large language models (LLMs) are built from structured high-dimensional objects such as token representations, weights, adaptation updates, caches, and activations, whose multilinear structure is underexploited by the conventional matrix-centric view. Tensor decompositions and tensor networks provide a principled algebraic language for this structure, yet the literature often treats them as isolated compression mechanisms. This survey organizes tensor methods for LLMs through two complementary views: a seven-stage lifecycle taxonomy covering tokenization, embeddings, pre-training, adaptation, compression, inference, and interpretability, and a component view covering embeddings, attention, and feed-forward networks. We provide unified notation and theoretical foundations, analyze tensorization strategies for individual Transformer components, and compare methods at each lifecycle stage while making differences in evaluation protocols and model scales explicit. We further connect tensor methods to neighboring efficiency techniques and probabilistic tensor networks. Finally, we synthesize open challenges and introduce $ρ_{\rm gap}$, a metric for the compression-realization gap between theoretical memory reduction and measured system-level speedup. By treating tensorization as a common structural principle, the survey provides a structured entry point to tensorized language models and clarifies when parameter savings can plausibly translate into memory efficiency, computational efficiency, or interpretability. The GitHub page dedicated to this paper is accessible at \href{https://github.com/ma-tt-a/awesome-tensor-methods-for-llms}{this https URL}.