arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Softmax丢弃的信息:用于证据积累的质量感知注意力

What Softmax Throws Away: Mass-Aware Attention for Evidence Accumulation

Minwoo Yu, Young-guk Ha

arXiv 2607.22781首次发表:更新:

发表机构

Smart Computing Laboratory, Department of Computer Science & Engineering, Konkuk University(智能计算实验室,计算机科学与工程系,建国大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究探讨模型内部表示中预测相关结构信息保留问题,提出质量感知注意力(MAA),将标准L1归一化推广到Lp族,在多模型和数据集上提升未来链接AUC等,定位为通用归一化原则提高表示信息性。

AI 中文摘要

高任务性能并不能表明模型在其内部表示中是否保留了与预测相关的结构信息。例如,时间图模型可以实现较高的未来链接AUC,但从相同表示中恢复基本图统计信息仍然困难。我们发现标准注意力加权平均中存在差距的一个来源:当证据模式重复时,分子和分母以相同速率增长,不同积累证据量的输入可产生相同的总和。我们提出质量感知注意力(MAA),将标准L1归一化推广到Lp族。在重复情况下,MAA使分子和分母以不同速率缩放,在表示量级中保留有贡献输入的有效数量。它不增加监督、参数、隐藏维度或显式计数特征,且在p = 1时恢复为标准注意力。在四个连续时间动态图模型和三个数据集上,MAA在12个模型 - 数据集单元中的11个中提高了未来链接AUC。从相同隐藏表示的线性恢复平均提高了4.49%,家族性校正后所有12个单元中的优先附着恢复均得到改善。在标记时间点过程、时间知识图、检索增强生成和时空点过程中也观察到一致证据。信息可访问性和任务效用仍然不同:MTPP中的负对数似然改善,TKG中的排名基本保留,RAG中的附加信息未改善诊断头,下游层归一化可消除STPP中的信号。这些结果将MAA定位为一种通用归一化原则,通过控制标准注意力中的重复不变性来提高面向预测器的表示信息性。

英文摘要

High task performance does not show whether a model retains prediction-relevant structural information in its internal representation. Temporal graph models, for example, can achieve high future-link AUC while basic graph statistics remain difficult to recover from the same representation. We identify one source of this gap in the weighted averaging used by standard attention: when an evidence pattern is repeated, the numerator and denominator grow at the same rate, so inputs with different amounts of accumulated evidence can produce the same aggregate. We propose Mass-Aware Attention (MAA), which generalizes standard L1 normalization to an Lp family. Under repetition, MAA makes the numerator and denominator scale at different rates, retaining the effective number of contributing inputs in the representation magnitude. It adds no supervision, parameters, hidden dimensions, or explicit count features, and recovers standard attention at p=1. Across four continuous-time dynamic graph models and three datasets, MAA improves future-link AUC in 11 of 12 model-dataset cells. Linear recovery from the same hidden representation increases by 4.49% on average, and preferential-attachment recovery improves in all 12 cells after family-wise correction. We also observe consistent evidence in marked temporal point processes, temporal knowledge graphs, retrieval-augmented generation, and spatio-temporal point processes. Information accessibility and task utility remain distinct: NLL improves in MTPP, ranking is largely preserved in TKG, additional information in RAG does not improve the diagnostic head, and downstream LayerNorm can erase the signal in STPP. These results position MAA as a general normalization principle for improving predictor-facing representation informativeness by controlling repetition invariance in standard attention.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑