发表机构
IBM Automation and AI; Hother; University of Warsaw; Centre for Credible AI, Warsaw University of Technology(IBM自动化与人工智能; 霍瑟; 华沙大学; 华沙理工大学可信人工智能中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究预训练语言模型嵌入的几何结构,提出黎曼均值池化方法,通过提取回拉度量并聚合,在三个数据集上RMP优于欧几里得均值池化,消融实验定位增益来源,训练好的编码器在特定数据集贡献额外信号。
AI 中文摘要
理解预训练语言模型嵌入的几何结构对可解释性和安全性很重要。我们研究句子级分类信号是否存在于上下文词元嵌入的黎曼几何中,并通过从学习到的编码器的解析雅可比矩阵中提取每个词元的回拉度量,并在对称正定(SPD)流形上用弗雷歇均值聚合它们来进行探究,我们将此过程称为黎曼均值池化(RMP)。在三个具有非平凡语言结构的数据集(CoLA、CREAK、RTE)上,RMP优于欧几里得均值池化,而在为消除注释驱动的词汇工件而构建的基准FEVER - Symmetric上,该方法正确地保持在随机水平。消融实验表明,随机初始化的编码器与弗雷歇聚合相结合,在三个有信号的数据集中的两个上已经超过了欧几里得池化,将增益来源定位到几何聚合而不是学习到的流形结构;训练好的编码器在CREAK(三个有信号的数据集中知识量最大的)上特别贡献了额外信号。
英文摘要
Understanding the geometric structure of pre-trained language model embeddings matters for interpretability and safety. We ask whether sentence-level classification signal lives in the Riemannian geometry of contextual token embeddings, and probe it by extracting per-token pullback metrics from a learned encoder's analytical Jacobian and aggregating them with the Fréchet mean on the symmetric positive definite (SPD) manifold; we call this procedure Riemannian Mean Pooling (RMP). Across three datasets with non-trivial linguistic structure (CoLA, CREAK, RTE), RMP outperforms Euclidean mean pooling, while on FEVER-Symmetric, a benchmark constructed to remove annotation-driven lexical artifacts, the method correctly stays at chance. Ablations show that a randomly initialised encoder combined with Fréchet aggregation already beats Euclidean pooling on two of the three signal-bearing datasets, localising the source of the gain to the geometric aggregation rather than to learned manifold structure; the trained encoder contributes additional signal specifically on CREAK, the most knowledge-heavy of the three signal-bearing datasets.