arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20912cs.LGquant-ph

量子模型能否像LLM一样扩展?

Do Quantum Models Scale Like LLMs?

  • Queen Mary University of London(伦敦玛丽女王大学)
  • National Taiwan University(国立台湾大学)
  • University of Waterloo(滑铁卢大学)
  • Perimeter Institute for Theoretical Physics(圆周理论物理研究所)

机构由 AI 辅助整理,请以论文原文为准。

David S. Berman, Ying-Jer Kao, Roger G. Melko, Alexander G. Stapleton

中文总结 AI 辅助

本研究通过RydbergGPT模型研究量子数据上的神经标度律,发现近临界统计类似自然语言,支持多尺度依赖性对稳定神经标度的作用。

中文摘要 AI 辅助

在这项工作中,我们研究了RydbergGPT的神经标度律,RydbergGPT是一种自回归Transformer模型,训练数据来自相互作用的里德伯原子阵列收集的量子比特投影测量数据。已知该量子系统在激光失谐参数变化时表现出临界点的有限尺寸残余。我们发现,在临界点附近,作为训练数据集大小函数的Transformer损失很好地由带有损失下限修正的幂律描述。然而,远离临界点时,幂律描述的质量显著降低。然后,我们使用熵归一化、有限样本校正的互信息“两点”函数比较了里德伯测量和自然语言语料的统计结构。我们发现,近临界的两点函数统计最接近自然语言中观察到的统计,而远离临界点的其他量子比特配置的两点函数衰减更快。这支持了多尺度依赖性有助于稳定神经标度,且标度行为应被视为模型-数据对属性的假设。

英文摘要

In this work, we study the neural scaling laws of RydbergGPT, an autoregressive transformer model trained on qubit projective measurement data gathered from interacting Rydberg atom arrays. The quantum system is known to exhibit a finite-size remnant of a critical point as the laser detuning parameter is varied. We find that near the critical point the transformer loss as a function of training dataset size is well described by a power-law with a loss floor correction. However, away from criticality the quality of the power-law description is substantially reduced. We then compare the statistical structure of both Rydberg measurements and natural-language corpora using an entropy-normalised, finite sample corrected mutual information "two-point" function. We find that near-critical statistics of the two point functions are closest to those observed in natural-language, whilst other qubit configurations far from the critical point have two-point functions that decay more rapidly. This supports the hypothesis that multi-scale dependence contributes to stable neural scaling, and that scaling behaviour should be viewed as a property of the model-data pair.

补充信息

↑