用于测量Transformer语言模型中语境个体化的工具包技术手册
Technical Manual for a Toolkit for Measuring Contextual Individuation in Transformer Language Models
- University of Campinas (Unicamp)(坎皮纳斯大学)
- Center for Electric Mobility Research (CEMOBE)(电动 mobility 研究中心(CEMOBE))
- Power Electronics Laboratories (LEPO)(电力电子实验室(LEPO))
- Institute of Computing (IC)(计算研究所(IC))
- Center for Logic, Epistemology and History of Science (CLE)(逻辑、认识论与科学史研究中心(CLE))
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本手册介绍了用于测量Transformer语言模型语境个体化的开源工具包,详述其流程与设计选择,为相关研究提供方法学与实现参考。
AI中文摘要:
Transformer语言模型在其嵌入层为一个词类型分配单一的、与语境无关的向量,但普遍认为其会在后续层中根据语境个体化该词的不同出现形式。要干净地验证这一观点,需要一种构造,该构造能固定词形,同时以受控、带标签的方式改变其语境和意图含义。本手册介绍了围绕这种构造构建的开源工具包,我们将其称为桥接形式(bridge form):即单个书面词,在两个或更多主题领域中重复出现且形式不变,但每个领域中含义不同。我们描述并论证了该工具包的每一个流程阶段:桥接形式及其源领域的声明式规范、从Wikipedia获取语料库、出现位置定位、分层表示提取、模型表示空间中分离度的领域配对轮廓测量,以及配对可视化协议。每个设计选择都与其旨在避免的方法失效模式一同呈现,这些失效模式包括:来自过于宽泛类别标签的含义污染、轮廓系数的多组偏差、子词分词错位,以及降维图中的轴可比性伪影等。本手稿是方法学和实现参考:它不报告或解释在任何特定模型或桥接形式集上运行该工具包的经验结果。该工具包、其完整源代码以及用于运行它的语料库均单独存档(第9节),并具有持久标识符,旨在被使用它产生和解释经验结果的研究作为工具引用。
英文摘要:
A transformer language model assigns a single, context-independent vector to a word type at its embedding layer, yet is widely believed to individuate that word's occurrences by context in its later layers. Testing this belief cleanly requires a construct that holds the word form fixed while its context and intended sense vary in a controlled, labeled way. This manual documents an open toolkit built around such a construct, which we call a bridge form: a single written word that recurs, unchanged, across two or more subject domains with a different sense in each. We describe, and justify, every stage of the pipeline: the declarative specification of bridge forms and their source domains, corpus acquisition from Wikipedia, occurrence localization, layer-wise representation extraction, a domain-pairwise silhouette measurement of separation in the model's representation space, and a paired visualization protocol. Each design choice is presented together with the methodological failure mode it is meant to avoid (sense contamination from overly broad category labels, the multi-group bias of the silhouette coefficient, subword-tokenization misalignment, and axis-comparability artifacts in dimensionality-reduced plots, among others). This manuscript is a methodological and implementation reference: it does not report or interpret empirical outcomes of running the toolkit on any particular model or bridge-form set. The toolkit, its full source, and the corpora used to exercise it are archived separately (Section 9) under a persistent identifier, and are intended to be cited as an instrument by studies that use it to produce and interpret empirical results.