从表达能力到样本复杂度:通过C-RASP为Transformer构建窄教师模型
From Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASP
浏览论文内容
中文总结 AI 辅助
研究Transformer的理论理解,通过提出手工权重等分析其表达能力。针对其可学习性研究少的问题,受损失景观分析启发,为用Transformer学习C-RASP结构提出初步样本复杂度界限。
中文摘要 AI 辅助
对Transformer的理论理解对于更好地理解大语言模型的能力和局限性至关重要。许多工作分析了基于注意力模型的表达能力。过去大量理论工作通过提出手工权重或使用计算复杂度论证来刻画哪些任务属于Transformer模型的假设类。然而,很少有工作研究此类解决方案的可学习性。在本工作中,受近期损失景观分析工作启发,我们朝着这一目标取得了进展,为用Transformer学习C-RASP结构提出了初步的样本复杂度界限。
英文摘要
A theoretical understanding of Transformers is crucial to better understand the capacities and limitations of large language models (LLMs). There is much work analyzing the expressivity of attention-based models. By proposing handcrafted weights or using computational complexity arguments, a large amount of past theoretical works have sought to characterize which tasks are and which are not in the hypothesis class of Transformer models. However, little work investigates the learnability of such solutions. In this work, we make progress towards this goal. Inspired by recent loss landscape analysis work, we propose preliminary sample complexity bounds for learning C-RASP constructions with Transformers.
发表机构
- Mila & Université de Montréal(米拉与蒙特利尔大学)
- University of Oxford(牛津大学)
- Saarland University(萨尔兰大学)
机构由 AI 辅助整理,请以论文原文为准。