arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FourierQK:滤波器形状、容许性与泄漏-覆盖定律

FourierQK: Filter Shape, Admissibility and the Leakage-Coverage Law

Athanasios Zeris

arXiv 2610.00009首次发表:更新:

发表机构

Independent Researcher(独立研究者)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过受控消融实验探究FourierQK中滤波器形状的影响,发现带通结构、容许性和频谱覆盖是关键,最优带宽带来显著增益,并提出了泄漏-覆盖定律,为双向和自回归注意力提供了设计指导。

AI 中文摘要

频率坍缩注意力(Frequency-collapse attention)[Zeris, 2026e]通过将Q/K点积替换为在习得频率下的带通滤波内积,相较于标准点积注意力取得了显著提升。一个自然的后续问题是:哪种滤波器形状效果最佳,以及为什么?我们针对五个关于滤波器属性的假设——直流抑制、奈奎斯特抑制、带宽、中心频率和多尺度覆盖——在字符级语言建模(TinyShakespeare,6层GPT)上进行了受控消融实验。我们的主要发现如下:(1)直流和奈奎斯特分量具有主动危害性(验证损失约等于2.0,相当于相位随机化),证实了振荡带通结构是必不可少的,而不仅仅是任何低维频谱摘要;(2)最优单尺度带宽为sigma约等于2个bin,以段落尺度(约70个token)为中心,相较于BASE-DOT实现了Delta = +1.15 nats的净增益;(3)容许滤波器(零均值,墨西哥帽DOG m=2)在相同尺度下优于非容许高斯滤波器,并对双边FFT泄漏提供部分保护;(4)双边FFT泄漏随频谱覆盖范围单调增加——窄带滤波器(间隙>+4)是干净的,宽带滤波器(间隙<+2)则泄漏严重;(5)字符尺度下的因果时域Morlet滤波器无法超越BASE-DOT(K=128个抽头覆盖了T=256上下文的50%),这促使在配套的MorletQK论文[Zeris, 2026f]中进行词级实验。综合发现(1)-(5),FourierQK被表征为在双向注意力设置(编码器风格,如BERT)中有效,其中在训练和推理时均可获得全序列上下文;自回归生成需要因果频谱变体,如MorletQK [Zeris, 2026f](解码器风格,如GPT)。代码可在以下网址获取:this https URL

英文摘要

Frequency-collapse attention [Zeris, 2026e] achieves large gains over standard dot-product attention by replacing the Q/K dot product with a bandpass-filtered inner product at a learned frequency. A natural follow-up question is: which filter shape works best, and why? We test five hypotheses about filter properties -- DC suppression, Nyquist suppression, bandwidth, centre frequency, and multi-scale coverage -- using a controlled ablation on character-level language modelling (TinyShakespeare, 6-layer GPT). Our main findings are: (1) DC and Nyquist components are actively harmful (val ~= 2.0, equivalent to phase randomisation), confirming that oscillatory bandpass structure is essential, not just any low-dimensional spectral summary; (2) the optimal single-scale bandwidth is sigma ~= 2 bins centred at paragraph scale (~70 tokens), giving a clean gain of Delta = +1.15 nats over BASE-DOT; (3) admissible filters (zero-mean, Mexican Hat DOG m = 2) outperform non-admissible Gaussians at the same scale and provide partial protection against bilateral FFT leakage; (4) bilateral FFT leakage scales monotonically with spectral coverage -- narrowband filters (gap > +4) are clean, wideband filters (gap < +2) are leaky; and (5) causal time-domain Morlet at character scale cannot beat BASE-DOT (K=128 taps covers 50% of T=256 context), motivating word-level experiments in the companion MorletQK paper [Zeris, 2026f]. Together, findings (1)-(5) characterise FourierQK as effective in bidirectional attention settings (encoder-style, e.g. BERT), where full-sequence context is available at both training and inference time; autoregressive generation requires a causal spectral variant such as MorletQK [Zeris, 2026f] (decoder-style, e.g. GPT). Code available at: https://github.com/AthanasiosZeris/energy-gated-attention

Comments9 pages, 1 figure, 2 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑