arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

句法与语义:Transformer如何学习深层依赖关系

Syntax vs. Semantics: How Transformers Learn Deep Dependencies

Jiangrui Zhao, Xiaoting Du

arXiv 2608.26139首次发表:更新:

AI 中文总结

该研究提出机械框架揭示Transformer学习深层语义依赖的梯度饥饿现象,解释CoT有效性,验证于多规模模型,提出拓扑对齐对比目标在变量绑定任务提升超2倍。

AI 中文摘要

大型语言模型展现出卓越的句法流畅性,但支配其获取深层语义依赖的优化动力学仍鲜为人知。我们提出一种机械框架,将该学习过程建模为表面统计与深层语义之间的竞争。理论分析发现了「梯度饥饿」现象:稀疏语义依赖的误差信号在早期优化阶段被主动抑制,这种抑制阻碍了结构推理的学习,使其表现为突然的相变。此外,该框架为思维链(Chain-of-Thought, CoT)策略的有效性提供了机械基础:通过将中间推理步骤外部化为具体 token,CoT 有效规避了隐式推理固有的抑制机制。我们在从小型 Transformer 到生产级模型(Llama-3.1-8B、Qwen2.5-Coder-7B)的不同规模上验证了这些发现。最后,基于该理论,我们提出一种拓扑对齐的对比目标,可显式修正梯度几何。在变量绑定任务上的实验表明,我们的方法实现的提升比标准交叉熵微调大2倍以上。

英文摘要

Large Language Models demonstrate remarkable syntactic fluency, yet the optimization dynamics governing their acquisition of deep semantic dependencies remain poorly understood. We propose a mechanistic framework that models this learning process as a competition between Surface Statistics and Deep Semantics. Our theoretical analysis identifies a ``Gradient Starvation" phenomenon where the error signals for sparse semantic dependencies are actively suppressed during early optimization. This suppression impedes the learning of structural reasoning and causes its emergence to manifest as a sudden phase transition. Furthermore, this framework offers a mechanistic basis for the effectiveness of Chain-of-Thought (CoT) strategies. By externalizing intermediate reasoning steps into concrete tokens, CoT effectively bypasses the suppression regime inherent to implicit reasoning. We validate these findings across scales ranging from toy transformers to production models (Llama-3.1-8B, Qwen2.5-Coder-7B). Finally, guided by this theory, we propose a topology-aligned contrastive objective that explicitly rectifies the gradient geometry. Experiments on variable binding tasks demonstrate that our method achieves an improvement that is over 2x larger than that obtained via standard cross-entropy fine-tuning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑