当余弦相似度无法反映对话模型中线性可访问的结构时
When Cosine Similarity Fails to Reflect Linearly Accessible Structure in Dialogue Models
AI总结:
本研究揭示在对话模型中,余弦相似度可能低估线性可访问的个性结构,通过监督子空间可弥补差距,但部分模型违反不变性标准。
AI中文摘要:
余弦相似度被广泛用于分析Transformer表示,其隐含假设是相似度能反映与任务相关的结构。我们研究了在对话条件下的大语言模型中,这一假设何时失效。在三个7-8B规模的聊天调优模型中,环境余弦相似度在相同的隐藏状态上显著低估了线性可解码的个性结构;具体而言,在30类任务上,线性探针AUC在0.73-0.97范围内,而余弦kNN在0.56-0.77范围内。一个低维监督子空间弥补了大部分差距,而匹配秩的PCA子空间则未能做到,并且在某些情况下性能下降。这种不匹配是依赖于情境的:在单句情感分类(SST-5)中不存在,并且匹配基数的对照实验排除了属性基数作为混淆因素。这种差距不会随对话轮次系统性增加,且任务对齐的子空间随时间保持稳定。然而,三个模型中有两个违反了预注册的子空间内可分离性不变性标准(|Delta AUC| <= 0.03),一个模型违反了预注册的轮次不变性标准(|Delta L| <= 0.05)。这些结果表明,即使对话表示中的任务对齐结构是线性可访问的,余弦相似度也可能无法反映该结构。
英文摘要:
Cosine similarity is widely used to analyze transformer representations, implicitly assuming that similarity reflects task-relevant structure. We study when this assumption fails in dialogue-conditioned large language models. Across three 7-8B chat-tuned models, ambient cosine similarity substantially underestimates linearly decodable persona structure on the same hidden states; numerically, linear probe AUC is in the 0.73-0.97 range while cosine kNN is in the 0.56-0.77 range on a 30-class task. A low-dimensional supervised subspace recovers much of this gap, whereas a matched-rank PCA subspace does not and in some cases degrades performance. This mismatch is regime-dependent: it is absent in single-sentence sentiment classification (SST-5), and a matched-cardinality control rules out attribute cardinality as a confound. The gap does not systematically increase across dialogue turns, and the task-aligned subspace remains stable over time. However, two of three models violate a pre-registered within-subspace separability invariance criterion (|Delta AUC| <= 0.03), and one model violates a pre-registered turn-invariance criterion (|Delta L| <= 0.05). These results show that cosine similarity can fail to reflect task-aligned structure in dialogue representations even when that structure is linearly accessible.