发表机构
University of Tübingen; German Center for Mental Health (DZPG)(蒂宾根大学; 德国心理健康中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多结果学习中共享注意力向量不稳定并坍缩的问题,提出结果索引注意力矩阵,经三个合成实验验证其能收敛到有意义表示,改善梯度型注意力学习。
AI 中文摘要
学习中的维度注意力通常实现为全局共享的注意力向量,其中每个刺激维度对应一个标量。这些标量由模型通过误差上的梯度下降进行学习,预测性特征获得更高的显著性。我们表明,在多结果学习(即模型预测多个结果)的情况下,该共享向量变得不稳定;它会坍缩到其边界,并阻止模型学习有意义的注意力调优,从而影响学习与泛化。为解决此问题,我们引入了一种基于结果索引的注意力矩阵,将全局共享的注意力调优转换为基于结果索引的表示。我们分析了不稳定的共享向量,并推导出其成立的条件。在实证方面,三个合成实验对所提出的注意力矩阵进行了基准测试,结果表明它们能收敛到有意义的表示,而共享注意力向量则无法做到。这些结果提示,基于结果索引的注意力矩阵是梯度型注意力过程的一种通用修复方法,可改进多结果条件下的学习模型。
英文摘要
Dimensional attention in learning is often implemented as a globally shared attention vector, where each stimulus dimension corresponds to a single scalar. These scalars are learned by models through gradient-descent on error, where predictive features acquire more salience. We show that under multi-outcome learning, where models predict more than one outcome, this shared vector becomes unstable; it collapses to its bounds and prevents the models from learning meaningful attentional tunings for learning and generalization. We address this by introducing an outcome-indexed attentional matrix that converts globally shared attentional tuning into an outcome-indexed representation. We present an analysis of the unstable shared vectors and derive the conditions under which it holds. Empirically, three synthetic experiments benchmark the proposed attention matrices and show that they converge to meaningful representations, something shared attention vectors fail to do. These results suggest that outcome-indexed attentional matrices are a general fix for gradient-based attentional processes, which improves models of learning under multi-outcome conditions.
Comments2 figures, 8 pages