复值相位相干Transformer
Complex-valued Phase-Coherent Transformers
浏览论文内容
中文总结 AI 辅助
本文提出相位相干Transformer,通过L2归一化查询和键的缩放余弦分数改进复值注意力,在多项任务上达到或超越实值基线,并首次以复值网络解决Path-X。
中文摘要 AI 辅助
复值Transformer继承了在原始复内积上的softmax注意力机制。在非原生复值领域之外,这种标准形式的表现接近随机水平,且此前没有复注意力机制被证明能够纠正这一问题。我们证明匹配必须是缩放余弦分数:对查询和键进行L2归一化,使分数反映它们的余弦相似度并忽略其幅度,并将该分数保持在一阶尺度。通过这种方式,相同的模型在两种不同门控下训练四个诊断任务;未归一化时,模型在ListOps和Needle上两种门控下均保持随机水平,在另外两个任务上则远低于基线,而归一化分数若置于过小的尺度也会失败。由此产生的相位相干Transformer(\PCT)系列在长程记忆、位置检索、层级推理、频域分类和物理复信号方面达到或超过最强的实值基线;在深度达20时无性能退化;其损失在61倍参数范围内呈对数线性下降。该系列的一个成员——复筛选结合相位相干递归——是首个真正解决Path-X的复值神经网络,其91.6%的可训练参数为复值,而S4为38.2%。我们将这些记录为复值神经网络中前所未有的泛化迹象。
英文摘要
Complex-valued Transformers have inherited softmax attention over the raw complex inner product. Outside natively complex domains this standard form stays near chance, and no complex attention had been shown to correct it. We show that the match must be a scaled cosine score: L2-normalise queries and keys, so the score reads their cosine similarity and ignores their magnitudes, and hold that score at order-one scale. With this the same models train on four diagnostic tasks under two different gates; without the normalisation they stay at chance on ListOps and Needle under both gates and fall far below on the other two, and a normalised score placed at too small a scale fails as well. The resulting family of phase-coherent Transformers (\PCT) matches or exceeds the strongest real-valued baseline across long-range memory, positional retrieval, hierarchical reasoning, frequency-domain classification and physical complex signals; it shows no degradation up to depth 20; and its loss decreases log-linearly over a 61-fold range of parameters. A member of the family, complex screening combined with a phase-coherent recurrence, is the first genuinely complex-valued neural network to solve Path-X, with 91.6% of its trainable parameters complex-valued against 38.2% for S4. We record these as signs of generalisation not previously seen in complex-valued neural networks.