发表机构
Université de Montréal; Mila, Quebec AI Institute; McGill University(蒙特利尔大学; 米拉魁北克人工智能研究所; 麦吉尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
HyperQ通过向冻结的扩散语言模型添加词元条件量子残差分支,以线性成本训练16-64量子比特电路,在基准测试中提升得分并减少微调数据需求。
AI 中文摘要
语言模型可以通过改变应用于单个词元的计算来适应。量子电路提供了这样一种方法,但在大型模型内部评估更宽的电路可能计算成本高昂。这里我们引入HyperQ,它为冻结的掩码扩散语言模型添加了词元条件的量子残差分支。量子残差分支是每个Transformer块中的一个模块,它读取词元的隐藏状态,发出该词元电路的坐标,执行该电路,并通过残差连接将测量值加回。骨干网络保持冻结,仅训练添加的分支。在每个分支内,一个轻量级电路超网络在共享的稀疏电路结构中发出词元特定的旋转角度、耦合强度和测量轴。所需的期望值具有精确的经典表达式,其评估成本随量子比特数线性增长,使得能够在1.1亿参数骨干网络内训练16至64量子比特的电路。在下游基准测试中,增加电路宽度将平均得分从47.65提高到54.30。在64量子比特时,HyperQ分别超过骨干网络及其低秩适应对应物4.71分和3.67分。HyperQ在20,000个提示-响应对上进行微调,而经典基线使用200,000个。这些发现支持词元条件电路发射作为量子增强语言建模的一种可行架构方法。
英文摘要
Language models can be adapted by changing the computations applied to individual tokens. Quantum circuits offer one such approach, but evaluating wider circuits inside a large model can be computationally demanding. Here we introduce HyperQ, which adds token-conditioned quantum residual branches to a frozen masked-diffusion language model. A quantum residual branch is a module in each transformer block that reads a token's hidden state, emits the coordinates of that token's circuit, executes it, and adds the measured values back through a residual connection. The backbone remains frozen, and only the added branches are trained. Within each branch, a lightweight circuit hypernetwork emits token-specific rotation angles, coupling strengths, and measurement axes in a shared sparse circuit structure. The required expectation values have an exact classical expression whose evaluation cost grows linearly with the qubit count, enabling circuits from 16 to 64 qubits to be trained within a 1.1-billion-parameter backbone. Across downstream benchmarks, increasing circuit width raises the average score from 47.65 to 54.30. At 64 qubits, HyperQ exceeds the backbone and its low-rank-adapted counterpart by 4.71 and 3.67 points, respectively. HyperQ is fine-tuned on 20,000 prompt-response pairs, compared with 200,000 for the classical baselines. These findings support token-conditioned circuit emission as a tractable architectural approach to quantum-augmented language modelling.
CommentsWork in progress