将量子电路与经典大语言模型分离
Separating quantum circuits from classical LLMs
- IBM Research(IBM研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究证明了低深度量子计算与经典大语言模型架构在分布和功能上的无条件分离,为大语言模型时代量子优势的研究奠定了基础。
AI中文摘要:
现代大语言模型——Transformer和扩散语言模型(DLM)——围绕两个经典算法任务构建:预测与生成。我们在这两种场景下证明了低深度量子计算与对应受限资源经典语言模型架构之间的无条件分离。具体而言,我们展示了以下结果:1. 分布分离:我们给出一个可由QNC⁰电路(即由有界扇入门构成的常深度量子电路族)采样的分布,即便允许现代DLM所依赖的次线性思维链及输出令牌修正/重屏蔽事件,任何具有浅调度和去噪的常轮扩散语言模型(DLM)都无法在恒定距离内采样该分布。2. 功能分离:我们展示了一个可在∧∘QNC⁰[log log n](即O(log log n)深度的QNC⁰电路族,其中n为输入长度,后接一个经典AND门)中计算的函数,任何计算该函数的常深度仅解码器Transformer都必须规模庞大:其宽度需达到n^Ω(1)。综上,本研究开启了大语言模型时代量子优势的研究。
英文摘要:
Modern large language models - transformers and diffusion language models - are built around two canonical algorithmic tasks: prediction and generation. We prove unconditional separations between low-depth quantum computation and the corresponding bounded-resource classical language-model architectures in both regimes. Concretely, we exhibit the following: 1. Distributional separation. We give a distribution that is sampleable by $\textsf{QNC}^0$ circuits (i.e., a family of constant-depth quantum circuits consisting of bounded fan-in gates) that no constant-round diffusion language model ($\textsf{DLM}$) with shallow scheduling and denoising can sample within constant distance, even when allowed sublinear chain-of-thought and output-token revision/remasking events, the very features modern $\textsf{DLM}$s rely on. 2. Functional separation. We exhibit a function computable in $\land \circ \textsf{QNC}^0[\log\log n]$ (i.e., a family of O$(\log\log n)$-depth $\textsf{QNC}^0$ circuits, where $n$ is the input length, followed by a single classical $\mathsf{AND}$ gate) such that any constant-depth decoder-only transformer computing the function must be large: it would have to have width $n^{Ω(1)}$. Together, our work initiates the study of quantum advantage in the era of large language models.