arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12671cs.AIcs.CC

Transformer的表达能力研究

On the Expressive Power of Transformers

  • University of California Santa Cruz(加利福尼亚大学圣克鲁兹分校)
  • IBM Research(IBM研究院)

机构由 AI 辅助整理,请以论文原文为准。

Phokion Kolaitis, Rik Sengupta

AI总结:

本文基于电路复杂性的概念与方法,概述了界定Transformer表达能力的部分精选研究结果,助力校准其作为语言识别器的表达能力。

AI中文摘要:

多层Transformer是当今几乎所有大型语言模型(LLM)的核心组件。由于其普遍性和计算能力,越来越多的研究旨在将Transformer作为语言识别器的表达能力与理论计算机科学领域研究了数十年的标准计算模型进行精确校准。在这一努力中,电路复杂性已大体成为分析Transformer表达能力的“正确”计算复杂性分支;原因在于,通过注意力、精度等各类资源对Transformer进行参数化,可直接与由门类型、规模、深度等资源参数化的不同电路类别进行比较。本文概述了利用电路复杂性的概念与方法界定Transformer表达能力的部分精选研究结果。

英文摘要:

Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today. Because of their ubiquity and computational capability, there is a rapidly growing body of work that aims to precisely calibrate the expressive power of transformers as language recognizers by comparing them against standard models of computation studied for decades by the theoretical computer science community. In this endeavor, circuit complexity has by and large emerged as the "correct" branch of computational complexity to analyze the expressive power of transformers; the reason is that parameterizing transformers by the various resources they use, such as attention and precision, leads to direct comparisons with different classes of circuits parameterized by resources such as type of gates, size, and depth. Here, we present an overview of selected results that delineate the expressive power of transformers using concepts and methods from circuit complexity.

补充信息

↑