arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

池化有帮助,学习权重在上下文学习中反而有害:分解组注意力

Pooling Helps, Learned Weighting Hurts In-Context: Decomposing Group Attention

Michael Fore, James Mason Inder, Mrishika Nair, Praneetha Vaddamanu, Sharlina Keshava

arXiv 2610.01831首次发表:更新:

发表机构

Amazon(亚马逊)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究分解组注意力机制,发现均匀池化在多变量和上下文学习预测中普遍有益,而学习权重在上下文学习中常有害,仅统一第一块注意力即可改善所有测试配置。

AI 中文摘要

组注意力由时间序列预测模型Chronos-2引入,在固定的补丁索引处对组的变量进行注意力操作,并同时服务于多变量(MV)和上下文学习(ICL)预测。我们不是整体评估这种跨变量注意力设计,而是探究机制中哪一部分带来了收益,并检验其在MV和ICL两种场景中的适用性。通过在推理时编辑注意力矩阵$\alpha$,我们分离了注意力头所包含的两条路径:V/O,它投影组的加权摘要;以及Q/K,它决定权重。均匀池化(没有Q/K加权的V/O)在我们20个传感器网络配置中的18个上产生正面效果,而学习到的加权(Q/K)则按组类型分化:其对MV的贡献是正面或可忽略的,但在10个传感器网络ICL配置中实质性地损害了8个,其中4个甚至比单变量推理更差。通过隔离不同层的影响,我们发现仅在第一块中统一$\alpha$就能改善我们测试的所有ICL配置。

英文摘要

Group attention, introduced by the time series forecasting model Chronos-2, attends over the variates of a group at a fixed patch index and serves both multivariate (MV) and in-context learning (ICL) forecasting. Rather than evaluating this cross-variate attention design as a whole, we ask which part of the mechanism earns the benefit and probe its applicability to both MV and ICL regimes. By editing the attention matrix $α$ at inference we separate the two pathways a head comprises: V/O, which projects a weighted summary of the group, and Q/K, which decides the weights. Uniform pooling (V/O without any Q/K weighting) is positive on 18 of our 20 sensor-network configurations, while the learned weighting (Q/K) splits by group type: its contribution is positive or negligible for MV, but materially degrades 8 of the 10 sensor-network ICL configurations, leaving 4 of them worse than univariate inference. By isolating the impact of different layers, we find that uniforming $α$ in the first block alone improves every ICL configuration we test.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑