arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17789quant-phcs.AIcs.CVphysics.app-ph

QiT:用于视觉识别任务的量子启发式Transformer

QiT: Quantum-Inspired Transformer for Visual Recognition Task

Badri N. Patro, Vijay Agneeswaran

首次发表
浏览论文内容

中文总结 AI 辅助

QiT提出一种量子启发式Transformer,通过角度编码、周期特征自注意力和门控乘法模拟,在经典硬件上实现量子模型的结构优势,在ImageNet-1K上达到78.3%准确率,同时保持标准Transformer的复杂度。

中文摘要 AI 辅助

量子机器学习提供了一种引人注目的表示视角:角度编码状态存在于希尔伯特空间中,在该空间中周期性相似性和相互作用可以被自然地表达。然而,在视觉识别中实现这一视角仍然困难,因为当前的量子神经网络受到量子比特数量有限、电路模拟和测量成本高昂、噪声以及在中型含噪声量子设备上优化不稳定的限制。我们研究是否可以将量子模型中的有用结构思想实现为可扩展的经典Transformer操作。我们引入了QiT,一种用于视觉任务的量子启发式Transformer,包含三个组件:(i)角度启发编码,将图像令牌映射到学习到的三角希尔伯特空间特征,类似于基于量子旋转的状态编码;(ii)对这些周期性特征进行自注意力,诱导出近似于量子保真度核的经典余弦核;(iii)门控乘法模拟,一种可训练的经典替代,用于变分电路中的相互作用项。所有组件都是可微的张量操作,因此QiT既不声称量子计算也不声称量子加速,并保留了标准视觉Transformer的$\mathcal{O}(N^2D)$注意力复杂度。在图像分类基准测试中,QiT与匹配的经典Transformer具有竞争力,同时避免了小型模拟量子Transformer所观察到的严重运行时间成本。QiT-B在ImageNet-1K上达到78.3%的top-1准确率,具有45.7M参数和11.5 GFLOPs。这些结果使QiT成为隔离和评估视觉识别中量子启发的归纳偏置的可扩展基线。

英文摘要

Quantum machine learning offers a compelling representational perspective: angle-encoded states inhabit Hilbert spaces in which periodic similarities and interactions can be expressed naturally. Realizing this perspective for visual recognition remains difficult, however, because present quantum neural networks are constrained by limited qubit counts, costly circuit simulation and measurement, noise, and unstable optimization on noisy intermediate-scale quantum devices. We investigate whether useful structural ideas from quantum models can instead be realized as scalable classical Transformer operations. We introduce QiT, a Quantum-inspired Transformer for vision tasks with three components: (i) angle-inspired encoding that maps image tokens to learned trigonometric Hilbert-space features analogous to quantum rotation-based state encoding; (ii) self-attention over these periodic features, inducing a classical cosine kernel approximated to quantum fidelity kernels; and (iii) gated multiplicative emulation, a trainable classical surrogate for interaction terms found in variational circuits. All components are differentiable tensor operations, so QiT claims neither quantum computation nor quantum speedup and retains the $\mathcal{O}(N^2D)$ attention complexity of a standard Vision Transformer. Across image-classification benchmarks, QiT is competitive with a matched classical Transformer while avoiding the severe runtime cost observed for a small simulated quantum Transformer. QiT-B reaches 78.3\% ImageNet-1K top-1 accuracy with 45.7M parameters and 11.5 GFLOPs. These results position QiT as a scalable baseline for isolating and evaluating quantum-motivated inductive biases in visual recognition.

发表机构

  • Microsoft(微软)

机构由 AI 辅助整理,请以论文原文为准。

↑