arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17288cs.CL

Q-干涉:内存高效的相位感知量子启发注意力机制

Q-Interference: Memory-Efficient Phase-Aware Quantum-Inspired Attention

Emama Nahid, Tahmid Imtiaz Imu, Huayue Gu, Liran Ma, Zhipeng Cai, Honghui Xu

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出Q-Interference,一种内存高效的相位感知量子启发注意力机制,通过精确三角分解避免大型中间张量,适配GPT架构,在基准实验中训练稳定且内存优势显著。

中文摘要 AI 辅助

GPT注意力机制通过点积相似度衡量令牌兼容性,该机制简单、有效且内存高效,但未显式建模强令牌特征应相互增强还是抑制。我们提出Q-Interference,一种用于自回归语言建模的全经典量子启发注意力机制,它为每个查询和键特征增加了振幅和学习到的相位,得到的注意力分数具有相位感知性:相位对齐则建设性贡献,相位冲突则破坏性贡献。尽管Q-Interference产生了比单纯相似度更丰富的交互规则,但Q-Interference的朴素实现需要一个大型令牌-对-特征交互张量,导致内存密集且通常不实用。为解决此限制,我们提出一种精确三角分解,通过两次标准矩阵乘法计算相同分数,避免了大型中间张量的物化。Q-Interference可直接适配GPT中的Transformer块,模型架构其余部分和下一个令牌预测目标保持不变。在公共基准数据集和基线模型上的实验表明,该重新表述在受控GPT风格设置中训练稳定,且比朴素相位感知干涉注意力具有一致的内存优势。这些结果支持本研究的特定贡献:一种精确的内存高效重新表述,使相位感知干涉注意力在标准GPT流程中变得实用。

英文摘要

GPT attention measures token compatibility through dot-product similarity. This mechanism is simple, effective, and memory-efficient. But it does not explicitly model whether strong token features should reinforce or suppress one another. We introduce Q-Interference, a fully classical quantum-inspired attention mechanism for autoregressive language modeling that augments each query and key feature with an amplitude and a learned phase. The resulting attention score is phase-aware which aligned phases contribute constructively while conflicting phases contribute destructively. Although Q-Interference yields a richer interaction rule than similarity alone, a naive implementation of Q-Interference requires a large token-pair-feature interaction tensor, making it memory-intensive and often impractical. To address this limitation, we propose an exact trigonometric factorization that computes the same score using two standard matrix multiplications avoiding materialization of the large intermediate tensor. Q-Interference fits directly into a Transformer block in GPT and leaves the remainder of the model architecture and next-token prediction objective unchanged. Experiments on public benchmark datasets and baseline models show that the proposed reformulation trains stably in a controlled GPT-style setting and provides a consistent memory advantage over naive phase-aware interference attention. These results support the specific contribution of this work: an exact memory-efficient reformulation that makes phase-aware interference attention practical within a standard GPT pipeline.

发表机构

  • Kennesaw State University(肯尼索州立大学)
  • Miami University(迈阿密大学)
  • Georgia State University(佐治亚州立大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑