arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.18413cs.CL

大语言模型中的卷积

Convolution for Large Language Models

  • Peking University(北京大学)
  • Huawei Technologies(华为技术有限公司)
  • Tsinghua University(清华大学)

机构由 AI 辅助整理,请以论文原文为准。

Yuchuan Tian, Yingte Shu, Wei He, Shuo Zhang, Tianchen Zhao, Chao Xu, Xinghao Chen, Yunhe Wang, Hanting Chen, Yu Wang

AI总结:

研究大语言模型中,轻量级深度卷积能否在不大幅增模型大小的情况下提供局部归纳偏差。通过实验发现特定位置应用卷积效果最佳,采用特定残差深度卷积设计,在多个模型和数据预算下提升了下游基准平均准确率,支持其作为自注意力补充建模短程交互。

AI中文摘要:

大语言模型很大程度上依赖于Transformer,其中自注意力提供全局令牌交互,但未明确编码自然语言的局部性。我们研究轻量级深度卷积能否在不大幅增加模型大小的情况下提供这种局部归纳偏差。宏观层面的消融实验在Qwen3 Transformer块中的17个位置比较卷积,发现对注意力前的投影查询、键和值应用卷积时效果最佳。随后的微观层面研究倾向于内核大小为k = 3的残差深度卷积,无需额外归一化或激活。在Qwen3模型和多个预训练数据预算中,该设计提高了七个下游基准的平均准确率,同时增加的参数不到0.01%。表示层面的案例研究进一步表明,卷积使重复令牌ID对其直接上下文更敏感。这些结果支持深度卷积作为自注意力的轻量级补充来建模短程令牌交互。

英文摘要:

Large language models (LLMs) largely rely on Transformers, where self-attention provides global token interaction but does not explicitly encode the locality of natural language. We study whether lightweight depthwise convolutions can supply this local inductive bias without materially increasing model size. Our macro-level ablation compares convolution at 17 locations in a Qwen3 Transformer block and finds the best results when convolution is applied to the projected queries, keys, and values before attention. A subsequent micro-level study favors a residual depthwise convolution with kernel size $k=3$, without additional normalization or activation. Across Qwen3 models and several pre-training data budgets, this design improves the average accuracy on seven downstream benchmarks while adding less than $0.01\%$ parameters. A representation-level case study further suggests that the convolution makes repeated token IDs more sensitive to their immediate context. These results support depthwise convolution as a lightweight complement to self-attention for modeling short-range token interactions.

补充信息

↑