arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

变压器真的无所不能吗?论跨任务归纳偏差的兼容性

Can Transformers Really Do It All? On the Compatibility of Inductive Biases Across Tasks

Damien Teney, Liangze Jiang, Hemanth Saratchandran, Simon Lucey

arXiv 2607.17624首次发表:更新:

发表机构

Idiap Research Institute; EPFL; Adelaide University(Idiap研究所; 洛桑联邦理工学院; 阿德莱德大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究探讨变压器对给定任务是否最优,提出为数据集优化变压器架构的方法,在算法玩具任务和代码语言建模数据集上实验,发现标准变压器非局部最优,简单替代方案有优劣,暗示有改进架构可支持多种能力。

AI 中文摘要

变压器用途广泛,设计在多种应用中基本一致。但它们对任何给定任务或数据集是否最优?答案对推动人工智能发展可能至关重要。方法:提出为给定数据集优化变压器架构的方法,用在留出数据上学习的函数替换重要非线性函数,然后在其他数据集上训练评估任务对之间的兼容性。结果:在算法玩具任务中识别出学习速度、分布内和分布外泛化及跨种子稳定性有显著改进的新架构,但非常特定于任务;在代码和语言建模数据集上也有改进且跨数据集和领域转移更好。结论:标准变压器很少是架构空间中的局部最优,简单替代方案性能更好但牺牲通用性,可能有改进架构能更好支持多种能力。

英文摘要

Transformers are remarkably versatile and their design is largely consistent across a variety of applications. But are they optimal for any given task or dataset? The answer may be key for pushing AI beyond merely scaling current designs. *Method.* We present a method to optimize a transformer architecture for a given dataset, which we use as a tool to study optimal task-specific inductive biases. This method replaces the most important non-linearities (GeLUs,;softmax) with functions learned on held-out data. We then train the resulting architectures on other datasets, as a way to evaluate the compatibility between pairs of tasks. Findings. On algorithmic toy tasks, we identify new architectures with dramatic improvements in learning speed, in- and out-of-distribution generalization, and stability across seeds. The new designs prove very task-specific however, and indicate that these tasks require inductive biases very different from those of standard transformers. On code and language modeling datasets, we also find architectures with consistent, yet smaller improvements. These designs transfer much better across datasets and domains (English & computer code). Implications. Our results show that standard transformers are rarely a local optimum in the space of architectures. Simple alternatives can perform much better but sacrifice universality. This suggests that there may be room for improved architectures that better support multiple capabilities simultaneously, such as fluency and robust reasoning.

CommentsPublished as a conference paper at ICLR 2026. https://github.com/idiap/lm-afs

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑