arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19171cs.LG

Lévy注意力:用于连续时间注意力的单次预测不确定性

Lévy Attention: Single-Pass Predictive Uncertainty for Continuous-Time Attention

Sotirios P. Chatzis, Loukas Papadoulas

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出Lévy注意力算子,可在单次前向传播中以闭式形式输出连续时间序列预测的不确定性,在t-PatchGNN等任务上表现优于蒙特卡洛dropout,且计算高效。

中文摘要 AI 辅助

针对不规则采样时间序列的深度模型可在任意连续时间戳处回答查询,但未报告每个答案应被信任的程度。本文表明注意力层本身可填补这一空白:通过合适的随机公式,生成每个预测的前向传播过程还能以闭式形式、无额外成本地输出该预测的可信程度。我们提出Lévy注意力,这是一种交叉注意力算子,其输出是针对非齐次泊松随机测度的随机积分:查询-键兼容性在连续(时间×通道)索引空间上组装出强度,测度在该强度下散射原子,输出则是这些原子处插值值场的平均值。在期望下,该算子会退化为平滑余弦核注意力,因此它可替代softmax层并以精确梯度训练。泊松结构保留了softmax所丢弃的闭式信息:证据Λ_q(总兼容性质量)和分歧trΣ_V(q)(值的离散度)。一个精确方差恒等式将二者结合为σ̂(q)=√(trΣ_V(q)φ(Λ_q)),这是采样算子的均方根偏差,由确定性前向传播生成且无需训练头。实验表明,分歧携带有效信号,而证据因子在密集数据上无信息,在稀疏数据上则具有强信息。在t-PatchGNN上,该算子替换仅造成最多5.6%的准确率损失(与匹配对照组相比),在最稀疏数据集上无损失;免费的分歧信号在匹配的五组随机种子套件上优于20次蒙特卡洛 dropout,且σ̂缩放后的校准高斯分布的零样本CRPS优于50次采样器;拆分共形包装器在每个水平达到名义覆盖率,单次传播可在1.4秒内对3383名未见过的患者按信任度排序。

英文摘要

Deep models for irregularly-sampled time series answer queries at arbitrary continuous timestamps, yet report nothing about how far each answer should be trusted. We show the attention layer itself can close that gap: with the right stochastic formulation, the pass that makes each prediction also reports, in closed form and at no extra cost, how far it should be trusted. We introduce Lévy Attention, a cross-attention operator whose output is a stochastic integral against an inhomogeneous Poisson random measure: query-key compatibilities assemble an intensity over a continuous (time x channel) index space, the measure scatters atoms under it, and the output averages an interpolated value field at those atoms. In expectation it reduces to a mollified cosine-kernel attention, so it replaces a softmax layer and trains with exact gradients. What softmax discards, the Poisson construction preserves in closed form: the evidence $Λ_q$ (total compatibility mass) and the disagreement $\mathrm{tr}\,Σ_V(q)$ (value spread). An exact variance identity makes their combination $\hatσ(q)=\sqrt{\mathrm{tr}\,Σ_V(q)\,φ(Λ_q)}$ the root-mean-square deviation of the sampled operator, emitted by the deterministic pass with no trained head. Empirically, disagreement carries the signal, while the evidence factor swings from uninformative on dense data to strongly informative on sparse. On t-PatchGNN the operator swap costs at most 5.6% accuracy against a matched control and nothing on the sparsest dataset. The free disagreement signal improves on 20-pass MC dropout across matched five-seed suites, and $\hatσ$ scales a calibrated Gaussian whose zero-sample CRPS beats a fifty-draw sampler; a split-conformal wrapper reaches nominal coverage at every level, and one pass ranks 3,383 unseen patients by trust in 1.4 seconds.

发表机构

  • Cyprus University of Technology(塞浦路斯理工大学)
  • Ethical AI Novelties

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑