自发欺骗与指令欺骗中的不对称性
Asymmetries in Spontaneous and Instructed Deception
浏览论文内容
中文总结 AI 辅助
该研究针对Llama-3.1-70B-Instruct模型,对比指令与自发欺骗场景,发现二者存在方向分量及迁移不对称性,且分类器训练、引导向量推导的最佳标记位置不同。
中文摘要 AI 辅助
大型语言模型有时会在未被指令的情况下欺骗用户,但现有关于模型欺骗的研究多集中于指令欺骗。我们在Llama-3.1-70B-Instruct模型中研究了指令欺骗与自发(未被指令)欺骗的关系,通过方向几何学、跨设置分类器及跨设置引导对比两种欺骗场景。研究发现两种欺骗场景共享约0.5余弦值的方向分量,且在检测与因果关系的迁移上存在不对称性:基于自发数据训练的分类器在指令数据上表现优于反之,基于指令推导的方向在引导自发提示时效果优于反之;此外,推导引导向量的最佳标记位置与训练及应用分类器的最佳标记位置存在差异。
英文摘要
Large language models sometimes deceive users without being instructed to. However, much of the study on deception in models involves instructed deception. We investigated the relationship between instructed and spontaneous (uninstructed) deception in Llama-3.1-70B-Instruct. We compared these two deception settings through direction geometry, cross-setting classifiers, and cross-setting steering. We found the two deception settings share a component of direction (cosine of approximately 0.5) and an asymmetry in the transfer between settings regarding detection and causation. Spontaneous trained classifiers performed better on instructed data than vice versa, and instructed derived directions performed better at steering spontaneous prompts than vice versa. Likewise the best token position to derive steering vectors from differed from the best token position to train and apply classifiers.