arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不变原子:语言模型表示中局部语义几何的稀疏坐标

Invariant Atoms: Sparse Coordinates of Local Semantic Geometry in Language Model Representations

Muhammad Ahtesham, Xin Zhong

arXiv 2609.36451首次发表:更新:

发表机构

University of Nebraska Omaha(内布拉斯加大学奥马哈分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出不变原子假设,通过学习共享语义框架和稀疏坐标重建语义位移,实现语义-干扰分离,并在多个实验中验证了原子的稳定性、泛化性和因果效应。

AI 中文摘要

大型语言模型通常在措辞、风格和句法发生重大变化时仍能保持语义不变,而微小的语义编辑却会系统地改变其隐藏表示。这表明语义变化可能沿着反复出现的局部方向组织。我们提出不变原子假设:局部语义运动在保持意义不变的变换下,沿稳定方向具有偏好的稀疏坐标。我们学习一个共享的语义框架和稀疏坐标,以重建语义位移同时抑制干扰变化,其中锚点相关的对角调制调整原子强度,而无需样本特定的旋转。实验上,原子表现出强烈的语义-干扰分离、稀疏重建、可复现的方向以及对模型预测的因果效应。学习到的几何结构泛化到未见过的语义邻域和干扰族,而局部重新加权提高了语义选择性并保持一致的全局到局部结构。原子签名在模型修改下也保持稳定。这些发现支持可复用的不变方向作为语言模型中局部语义几何的稀疏坐标系。

英文摘要

Large language models often preserve meaning despite substantial changes in wording, style, and syntax, while small semantic edits can systematically alter their hidden representations. This suggests that semantic variation may be organized along recurring local directions. We propose the Invariant Atom Hypothesis: local semantic motion admits preferred sparse coordinates along directions that remain stable under meaning-preserving transformations. We learn a shared semantic frame and sparse coordinates that reconstruct semantic displacements while suppressing nuisance variation, with anchor-dependent diagonal modulation adjusting atom strengths without sample-specific rotations. Empirically, the atoms exhibit strong semantic--nuisance separation, sparse reconstruction, reproducible directions, and causal effects on model predictions. The learned geometry generalizes to unseen semantic neighborhoods and nuisance families, while local reweighting improves semantic selectivity and preserves a consistent global-to-local structure. Atom signatures also remain stable under model modification. These findings support reusable invariant directions as a sparse coordinate system for local semantic geometry in language models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑