发表机构
Université Laval; Institut intelligence et données; Mila – Québec Artificial Intelligence Institute(拉瓦尔大学; 智能与数据研究所; 米拉-魁北克人工智能研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对深度学习的组合泛化难题,提出了基于张量积表示(TPRs)的注意力机制,经实验验证其在组合泛化性能上优于现有架构组件,为系统性泛化模型提供了新方向。
AI 中文摘要
系统性泛化仍是深度学习中的重大挑战。尤其是组合泛化——即对已知变异因素的新配置进行泛化——对人类而言轻而易举,但对依赖统计关联而非显式结构表示的标准神经架构而言却十分困难。我们引入了一种新的架构组件,它将结构化归纳偏置嵌入深度学习:一种针对张量积表示(TPRs)运行的注意力机制。通过在组合任务上开展的受控实验,我们表明这种TPR-注意力机制在组合泛化方面优于现有的架构组件。这些结果凸显了将显式组合结构整合到神经注意力中的价值,并为具备系统性泛化能力的模型指明了一条有前景的路径。
英文摘要
Systematic generalization remains a significant challenge in deep learning. In particular, combinatorial generalization - generalizing to new configurations of known factors of variation - is effortless for humans but difficult for standard neural architectures that rely on statistical correlations rather than explicit structural representations. We introduce a new architectural component that embeds structured inductive bias into deep learning: an attention mechanism operating over tensor-product representations (TPRs). Through controlled experiments on compositional tasks, we show that this TPR-attention mechanism outperforms existing architectural components in combinatorial generalization. These results highlight the value of integrating explicit compositional structure into neural attention and point toward a promising path for models capable of systematic generalization.
Journal refICLR 2026 Workshop on Geometry-grounded Representation Learning and Generative Modeling