AI 中文总结
本研究提出将泛化探测作为子空间选择的方法,利用Llama-3.1-8B-Instruct的主成分实现分布外欺骗检测的跨域迁移,有效缩小了基准与最优方法的性能差距。
AI 中文摘要
线性探测可用于检测语言模型激活内部的行为与概念,但可能无法迁移到分布外示例。在研究Llama-3.1-8B-Instruct探测在3个保留的欺骗检测数据集上的泛化性能时,我们发现将输入投影到激活训练分布的一小部分主成分(PC)上,可实现跨域迁移,其性能几乎与直接在测试分布上训练的探测相当。此外,我们发现PC解释可用于找到这些可迁移PC的子集。通过使用大语言模型(LLM)评判对每个PC进行评分,判断其最活跃/最不活跃的示例是否暗示可迁移的欺骗方向,随后对得分最高的PC进行探测,我们在内幕交易报告数据集上缩小了基线与最优方法(oracle)之间78%的差距,在沙盒操作(Sandbagging)数据集上缩小了25%的差距。源探测权重较高的方向似乎编码了源特定的表面特征,而实际可迁移的方向则以更抽象的方式编码相同的对比,这种方式可通过自然语言描述捕捉。总体而言,我们的结果表明,探测的分布外鲁棒性在很大程度上由子空间选择决定。
英文摘要
Linear probes can be used to detect behaviors and concepts inside language model activations, but may fail to transfer to out-of-distribution examples. When studying the generalization performance of Llama-3.1-8B-Instruct probes over 3 held-out deception detection datasets, we find that projecting inputs onto a small subset of principal components (PCs) from the training distribution of activations enables cross-domain transfer that nearly matches the performance of probes trained directly on the test distribution. Furthermore, we find that PC interpretations can be used to find a subset of those transferable PCs. By using an LLM judge to score each PC on whether its most/ least activating examples imply a transferable deception direction, then probing on the highest-scoring PCs, we close the baseline-to-oracle gap by 78% on Insider Trading Report and by 25% on Sandbagging. The directions a source probe weights heavily appear to encode source-specific surface features, while the directions that actually transfer appear to encode the same contrast more abstractly, in a way natural language descriptions can capture. Broadly, our results suggest that the OOD robustness of probes is largely determined by subspace selection.