AI 中文总结
研究证明意义同一性并非嵌入几何属性,而是联合前向传播中的计算操作;独立编码器性能差,联合读取器可高精度识别并蒸馏,挑战了传统检索范式。
AI 中文摘要
意义同一性(即两个句子在措辞变化后是否表达相同意思)在检索和RAG中被视为独立编码的句子向量的几何事实。我们证明,对于冻结的现成编码器和语言模型,情况并非如此:同一性是在两个句子共享一次前向传播时计算出来的,而不是这些系统所携带的嵌入几何的属性。在重叠匹配的PAWS-X上,专用编码器(BGE、E5、GTE、MiniLM、E5-Mistral-7B)的英语确认AUC仅为0.55-0.65(密集峰值0.70)。Llama 3、Mistral和Qwen的独立编码最后标记状态表现并不更好;两个向量的晚期融合仍接近随机水平。同一探针在联合前向传播上达到0.90-0.96(从1.5B到32B),在伙伴洗牌下崩溃,处于中间深度,在3B时饱和接近0.94,在GPT-2 XL(0.76)中表现较弱。这一差距在Llama风格模型之外的其他因果语言模型、双向编码器(DeBERTa、RoBERTa)和编码器-解码器(Flan-T5、T5、BART)中依然存在。固定或线性读取器在冻结的独立编码上永远无法解锁同一性;非线性成对读取器仅在完整的49k对PAWS训练分割上恢复部分同一性(0.68-0.87)。现成的重排序器表现分化:BGE-reranker-large达到0.94,而MS-MARCO和Jina保持在0.55-0.64。独立训练的家庭计算相同的关系,一个1.5B的联合读取器可以从无标签的教师分数中蒸馏出它,而教师的独立向量的任何线性函数都无法做到。双编码器可以微调以适应PAWS(0.87-0.93),但迁移和STS-B受到影响。余弦比较的是措辞邻域;同一性是一个廉价的计算算子,而不是任一句子向量的属性。
英文摘要
Meaning identity (whether two sentences say the same thing after wording changes) is treated in retrieval and RAG as a geometric fact about independently encoded sentence vectors. We show that, for frozen off-the-shelf encoders and language models, it is not: identity is computed when both sentences share one forward pass, and is not a property of the embedding geometry those systems ship. On overlap-matched PAWS-X, purpose-built encoders (BGE, E5, GTE, MiniLM, E5-Mistral-7B) reach English confirm AUC only 0.55-0.65 (dense peak 0.70). Independently encoded last-token states of Llama 3, Mistral, and Qwen do no better; late fusion of the two vectors stays near chance. The same probe on a joint forward pass reaches 0.90-0.96 from 1.5B to 32B, collapses under partner shuffle, is mid-depth, saturates near 0.94 by 3B, and appears more weakly in GPT-2 XL (0.76). The gap holds beyond Llama-style models on other causal LMs, bidirectional encoders (DeBERTa, RoBERTa), and encoder-decoders (Flan-T5, T5, BART). Fixed or linear readers over frozen independent encodings never unlock identity; nonlinear pair readers recover part of it only on the full 49k-pair PAWS train split (0.68-0.87). Off-the-shelf rerankers split: BGE-reranker-large reaches 0.94, while MS-MARCO and Jina stay at 0.55-0.64. Independently trained families compute the same relation and a 1.5B joint reader can distill it from unlabelled teacher scores, while no linear function of the teachers own independent vectors can. Bi-encoders can be fine-tuned to fit PAWS (0.87-0.93), but transfer and STS-B suffer. Cosine compares wording neighbourhoods; identity is a cheap computed operator, not a property of either sentence vector.
Comments12 pages, 2 figures. Code to be released