你看不见的是你所学的:受限证据可见性有利于共享基因组语言模型社群中的组合泛化
What You Can't See Is What You Learn: Slot-Selective Evidence Masking Favors Compositional Generalization in Shared-Genome Language-Model Societies
AI总结:
该研究探究受限证据可见性对共享基因组语言模型社群的影响,发现受限可见性可显著提升组合泛化能力,虽未通过完整预注册测试,但利于形成可重用的值索引接口。
AI中文摘要:
多模块系统通常会让每个模块都接触到完整的输入。我们测试限制证据可见性是否会改变基于梯度的训练所发现的解决方案。四模块社群共享一个冻结的预训练语言模型和一个低秩适配器,仅通过固定中继中的两个模型宽度连续向量进行通信。在一项前瞻性密封的自然语言函数组合任务中,我们训练了十对匹配的受限/全局模型,它们共享初始化字节、训练顺序、标记布局、参数和计算;仅注意力掩码不同。在10对中的9对里,受限社群在两种深度下的表现都比其全局可见的孪生模型高出至少20个点,配对优势中位数分别为0.7648和0.6050。切断通信会使每个受限社群的表现降至随机水平,且在训练中从未出现过复合函数的程序上,深度三的优势仍为0.558。在6个经审核的受限社群中,相同值数据包移植在所有测试接口上的行为保持在0.94至1.00之间;破坏性干预会使性能崩溃;反事实数据包会将输出重定向至数学预测的答案。唯一表现优异的全局模型也需要通信,但其相同值数据包在不同回合间不可互换。因此,受限可见性并非组合所必需;在该协议下,它大幅提高了泛化中继的概率,并有利于可重用的、值索引的接口。不过,完整的预注册测试组正式失败,因为受限组的深度三准确率中位数为0.6988,低于0.70的下限。早期的资格队列同样产生0/10的完整通过:一个模型满足所有任务性能门槛,但所有10个模型都未能通过普通语言保留,将系统限制为明确任务门控的使用。
英文摘要:
Multi-module neural systems often expose every module to the full input. We test whether a slot-selective evidence-masking regime -- restricting each module to its own evidence span -- changes which solutions gradient-based training discovers. Four-cell societies share one frozen pretrained language model and one low-rank adapter, communicating only through two model-width continuous vectors in a fixed relay. On a prospectively sealed natural-language function-composition task, we train ten matched restricted/global pairs identical except for the attention mask. Restricted-visibility societies outperform their globally visible twins by at least 20 percentage points at both depths in 9 of 10 pairs, with median paired advantages of 0.7648 and 0.6050. Cutting communication reduces every restricted society to chance, and in a post hoc collision-stratified analysis the depth-three advantage remains 0.558 on programs whose complete affine map never appeared in training. In six post hoc-selected restricted societies, packet interventions on correctly answered held-out episodes are consistent with approximately value-indexed relay states; the sole high-performing global model also requires communication, but its same-value packets are not interchangeable across episodes. Thus restricted visibility is not necessary for composition. Under the tested seeds, streams, task world, and training budget, the masking regime strongly shifted which solutions training discovered: a post hoc mask crossover finds both arms mask-native. Because the restricted mask both blocks foreign evidence and implicitly identifies each cell's assigned slot, attribution to evidence visibility alone awaits a role-marked control. The preregistered battery nevertheless formally fails because restricted-arm median depth-three accuracy is 0.6988, below the 0.70 floor; an earlier qualification cohort yielded 0/10 complete passes.