arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

谁获得 Token,它携带什么?大型语言模型中不平等的名称支持与概念访问

Who Gets a Token, and What Does It Carry? Unequal Name Support and Concept Access in Large Language Models

Mir Tafseer Nayeem, Davood Rafiei

arXiv 2609.34065首次发表:更新:

发表机构

University of Alberta(阿尔伯塔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究揭示大型语言模型中名字的词汇支持不平等,提出NameTrace框架衡量概念可访问性,证明词汇差异影响任务表现,强调行为可比性始于词汇可比性。

AI 中文摘要

名字是个人标识符,但它们也承载社会意义,并被广泛用于评估语言模型如何对待不同的人。此类评估通常假设匹配的名字是可比较的模型输入。我们表明,这一假设在词汇接口层面常常失效:匹配的名字未必是匹配的输入。有些名字获得直接的单词级访问,而另一些则由多个子词组装而成,造成不平等的名称表面支持。在近五十万个名字和12个与LLM相关的分词器中,直接的词汇访问具有高度选择性、依赖模型,并在与种族和性别相关的名字元数据上分布不均。我们引入NameTrace,一个模型原生的、细粒度的、行为前的框架,用于衡量不平等的名称表面支持是仅停留在词汇属性层面,还是在任务相关的内部表征中变得可见。NameTrace通过模型自身在任务特定形容词轴上的概率(带有连续任务对齐权重)来衡量概念可访问性。在相同种族/民族-性别关联层内的匹配原子名和短碎片名上,支持度预测了在奖学金、招聘、临床评估和贷款等任务中概念可访问性的系统性差异。这些差异在所有八个匹配层中持续存在,跨越模型家族,并迁移到未见过的名字。隐藏状态干预进一步表明,所测量的任务方向具有下游影响力,能改变后续的受限选择。因此,不平等的词汇支持在输入层面具有人口统计学结构,并在任务相关的模型计算中保持可见。NameTrace使词汇可比性变得可测量,支持一个更广泛的原则:行为可比性始于词汇可比性。

英文摘要

Names are personal identifiers, but they also carry social meaning and are widely used to evaluate how language models treat different people. Such evaluations typically assume that matched names are comparable model inputs. We show that this assumption often fails at the lexical interface: matched names are not necessarily matched inputs. Some names receive direct single-token access, while others are assembled from multiple subwords, creating unequal name-surface support. Across nearly half a million first names and 12 LLM-associated tokenizers, direct lexical access is highly selective, model dependent, and uneven across race- and gender-associated name metadata. We introduce NameTrace, a model-native, fine-grained, pre-behavioral framework for measuring whether unequal name-surface support remains a vocabulary property or becomes visible in task-relevant internal representations. NameTrace measures concept accessibility from the model's own probabilities over task-specific adjective axes with continuous task-aligned weights. On matched atomic and short-fragmented names within the same race/ethnicity--gender-associated strata, support predicts systematic differences in concept accessibility across fellowship, hiring, clinical assessment, and lending. These differences persist across all eight matched strata, extend across model families, and transfer to unseen names. Hidden-state interventions further show that the measured task directions have downstream leverage, shifting later constrained choices. Unequal lexical support is therefore demographically structured at the input and remains visible in task-relevant model computation. NameTrace makes lexical comparability measurable, supporting a broader principle: behavioral comparability begins with lexical comparability.

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑