arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

变分伊辛注意力(VIA):适用于科学领域的定制注意力很重要

Variational-Ising-Attention:Tailored Attention Matters for Science

Rui Wang

arXiv 2607.23634首次发表:更新:

发表机构

DeepSeek-AI; Moonshot AI(深势科技人工智能; 月之暗面人工智能)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究针对科学任务中softmax注意力独立性假设的不足,提出变分伊辛注意力(VIA),通过伊辛模型增强softmax归一化,在逆合成反应中心预测任务中实例化VIA,实验证明其优于标准softmax注意力,强调科学问题需定制注意力。

AI 中文摘要

注意力通过带有softmax归一化的查询-键评分实现上下文建模。受工业长上下文需求驱动,主流研究趋向稀疏性和效率,但softmax的独立性假设依然存在。对于无长令牌约束的科学任务,更丰富的结构化耦合可能至关重要,因此定制注意力既可行又更合适。为此,我们提出变分伊辛注意力(VIA),用相互作用的伊辛模型增强softmax归一化。注意力模式通过变分平均场推理从可学习的成对耦合中出现,将注意力从孤立项目的排序重新定义为相互作用实体的集体状态。我们在逆合成反应中心预测任务上实例化VIA,该任务受协同断键约束支配。综合实验表明VIA始终显著优于标准softmax注意力。更广泛地说,我们的发现表明对于科学问题,最佳解决方案不是通用效率,而是与内在领域结构相匹配的适当定制注意力。这项工作提供了该范式的理论基础和实证验证实例。

英文摘要

Attention enables context modeling via query-key scoring with softmax normalization. Driven by industrial long-context demands, mainstream research has converged toward sparsity and efficiency, yet softmax's independence assumption persists. For scientific tasks unburdened by long-token constraints, however, richer structured coupling may often be essential, making tailored attention both viable and more appropriate. To this end, we propose Variational-Ising-Attention (VIA), which augments softmax normalization with an interacting Ising model; attention patterns emerge from learnable pairwise couplings via variational mean-field inference, extending attention from a ranking over isolated items to a collective state over interacting entities. We instantiate VIA on retrosynthesis reaction center prediction and, as a controlled internal ablation, on protein residue contact prediction, two structured prediction tasks governed by cooperative constraints: cooperative bond-breaking for retrosynthesis and inter-residue interactions for protein contact prediction. Comprehensive experiments across model variants, coupled with mechanistic analyses, demonstrate that VIA substantially outperforms standard softmax attention. More broadly, our findings suggest that for scientific problems, the optimal solution is not general-purpose efficiency, but appropriately tailored attention aligned with intrinsic domain structure. This work provides a theoretically grounded and empirically validated instantiation of this paradigm.

Comments24 pages, ~30 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑