逻辑的几何学:分层诱导语义结构与稳健推理
The Geometry of Logic: Stratification Induces Semantic Structure and Robust Reasoning
浏览论文内容
中文总结 AI 辅助
本研究提出STRAT架构,通过将残差流分为正交的数据与类型子空间,实现数据-控制分离,在算术任务中将中位数OOD误差降低35倍,并在11个数据集上平均准确率提升26个百分点,显著增强逻辑推理的稳健性。
中文摘要 AI 辅助
基于Transformer的语言模型在符号任务上表现良好,但仍不清楚它们是否学习到可泛化的规则,还是依赖于统计捷径。机制研究将算法行为与结构化内部表征联系起来,这支持了一个假设:稳健推理受益于将值与其控制操作的类型分离。将这种分离作为架构原语能否提高逻辑机制的可学习性和泛化能力?我们引入了STRAT(分层寄存器与类型),它将残差流划分为正交的数据子空间和类型子空间,并使用基于类型的注意力和门控来管理数据变换。受控的算术消融实验识别出与数据-控制干扰相关的三种失败模式:线性陷阱、梯度墙和开放门陷阱。机制分析揭示了可解释的逻辑结构,在算术任务中,STRAT将中位数OOD误差相对于Transformer基线降低了35倍。在跨越10个任务的11个数据集上,每个数据集使用10个基础示例训练,并在适用时使用相同的任务特定增强,STRAT在平均准确率上优于Transformer基线,平均高出26个百分点。在分布偏移下,STRAT的平均准确率仅下降2.39个百分点,而Transformer则下降11.75个百分点。
英文摘要
Transformer-based language models perform well on symbolic tasks, yet it remains unclear whether they learn generalizable rules or rely on statistical shortcuts. Mechanistic studies link algorithmic behavior to structured internal representations, motivating the hypothesis that robust reasoning benefits from separating values from the types that control their manipulation. Can making this separation an architectural primitive improve the learnability and generalization of logical mechanisms? We introduce \textbf{STRAT} (\textbf{ST}ratified \textbf{R}egisters \textbf{A}nd \textbf{T}ypes), which partitions the residual stream into orthogonal Data and Type subspaces and uses Type-based attention and gating to govern Data transformations. Controlled arithmetic ablations identify three failure modes associated with data-control interference: the Linear Trap, Gradient Wall, and Open Gate Trap. Mechanistic analysis reveals interpretable logical structure, and in arithmetic, STRAT reduces median OOD error 35-fold relative to a Transformer baseline. On each of 11 datasets spanning 10 tasks, STRAT outperforms the Transformer baseline in mean accuracy, by 26 percentage points on average, with both models trained from 10 base examples per dataset using identical task-specific augmentation where applicable. Under distribution shift, STRAT's mean accuracy drops by only 2.39 percentage points, compared with 11.75 for the Transformer.