arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.11760cs.LGcs.CL

从表达能力到样本复杂度:通过C-RASP为Transformer构建窄教师模型

From Expressivity to Sample Complexity: Narrow Teachers for Transformers via C-RASP

Michael Rizvi-Martel, Satwik Bhattamishra, Guillaume Rabusseau, Michael Hahn

首次发表
浏览论文内容

中文总结 AI 辅助

研究Transformer的理论理解,通过提出手工权重等分析其表达能力。针对其可学习性研究少的问题,受损失景观分析启发,为用Transformer学习C-RASP结构提出初步样本复杂度界限。

中文摘要 AI 辅助

对Transformer的理论理解对于更好地理解大语言模型的能力和局限性至关重要。许多工作分析了基于注意力模型的表达能力。过去大量理论工作通过提出手工权重或使用计算复杂度论证来刻画哪些任务属于Transformer模型的假设类。然而,很少有工作研究此类解决方案的可学习性。在本工作中,受近期损失景观分析工作启发,我们朝着这一目标取得了进展,为用Transformer学习C-RASP结构提出了初步的样本复杂度界限。

英文摘要

A theoretical understanding of Transformers is crucial to better understand the capacities and limitations of large language models (LLMs). There is much work analyzing the expressivity of attention-based models. By proposing handcrafted weights or using computational complexity arguments, a large amount of past theoretical works have sought to characterize which tasks are and which are not in the hypothesis class of Transformer models. However, little work investigates the learnability of such solutions. In this work, we make progress towards this goal. Inspired by recent loss landscape analysis work, we propose preliminary sample complexity bounds for learning C-RASP constructions with Transformers.

发表机构

  • Mila & Université de Montréal(米拉与蒙特利尔大学)
  • University of Oxford(牛津大学)
  • Saarland University(萨尔兰大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑