arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

关系注意力用于数据高效语言建模

Relational Attention for Data-Efficient Language Modeling

Adrian Brasoveanu, Ece Takmaz, Jakub Dotlačil

arXiv 2609.20530首次发表:更新:

发表机构

UC Santa Cruz; Utrecht University(加州大学圣克鲁兹分校; 乌得勒支大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出Relational BabyLM,结合双注意力Transformer和下一潜在预测目标,提升数据效率,在BabyLM 2026挑战赛严格轨道上排名第6,优于GPT-2基线。

AI 中文摘要

我们提出了Relational BabyLM,这是对BabyLM 2026挑战赛的一个系统提交,它在单个仅解码器Transformer中结合了两种认知动机的归纳偏置。在架构上,我们用双注意力Transformer(DAT)取代了标准自注意力,该Transformer将对象级(“感觉”)词汇特征的路径选择与结构/关系信息分离开来(Altabaa和Lafferty,2025;Altabaa等人,2024;Webb等人,2024;Kerg等人,2022;Webb等人,2021)。从自注意力中解耦的关系注意力(RA)在纯关系任务上大大提高了数据效率和训练样本外泛化能力,但语言建模需要对象级和关系信息既被整合又被解耦,而基于RA的语言模型在很大程度上仍未得到探索。BabyLM的数据受限训练和全面评估是测试这种数据效率是否能够迁移的理想试验场。作为训练干预,我们添加了一个下一潜在预测(NextLat;Teoh等人,2026)目标,该目标鼓励隐藏状态逐步将历史压缩为密集的信念状态。架构是结构语言泛化的主导因素;目标是次要但仍然显著的。在1000万词规模下,DAT的三种关系注意力类型(完整RA与更简单的RCA和DisRCA变体)在很大程度上可互换;完整RA在1亿词规模下领先。我们还引入了一种新颖的符号检索机制(基于RoPE,而非学习的相对符号),它在不增加参数的情况下匹配了学习符号库。在严格(1亿词)轨道上,我们最好的模型在撰写本文时总体排名55个中的第6位,在排行榜的NLP任务子集上排名55个中的第3位;我们的两个最强模型在大多数基准测试上优于GPT-2基线,其中一个在严格轨道参赛作品中获得了最高的EWoK分数。

英文摘要

We present Relational BabyLM, a system submission to the BabyLM 2026 challenge that combines two cognitively motivated inductive biases in a single decoder-only Transformer. Architecturally, we replace standard self-attention with a Dual Attention Transformer (DAT), which separates the routing of object-level ("sensory") lexical features from structural/relational information (Altabaa and Lafferty, 2025; Altabaa et al., 2024; Webb et al., 2024; Kerg et al., 2022; Webb et al., 2021). Relational attention (RA) disentangled from self-attention greatly increases data efficiency and out-of-training-sample generalization on purely relational tasks, but language modeling requires object-level and relational information to be integrated as well as disentangled, and RA-based LMs have remained largely unexplored. BabyLM's data-constrained training and comprehensive evaluation is an ideal testing ground for whether that data efficiency transfers. As a training intervention, we add a Next-Latent Prediction (NextLat; Teoh et al. 2026) objective that encourages hidden states to compress history incrementally into a dense belief state. Architecture is the dominant factor for structural linguistic generalization; the objective is secondary but still significant. DAT's three relational attention types (full RA vs. the simpler RCA and DisRCA variants) are largely interchangeable at 10M words; full RA pulls ahead at 100M. We also introduce a novel symbol-retrieval mechanism (RoPE-based, as opposed to learned, relative symbols) that matches learned symbol libraries while adding no parameters. On the strict (100M-word) track, our best model ranks 6th of 55 overall and 3rd of 55 on the leaderboard's NLP-task subset at the time of writing; our two strongest models outperform the GPT-2 baseline on most benchmarks, with one attaining the highest EWoK score among strict-track entries.

CommentsBabyLM Workshop, EMNLP 2026. Source code: https://github.com/abrsvn/babylm_dat_2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑