发表机构
Predictably Weird; Amazon(Predictably Weird; 亚马逊)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出高特异性的模型不一致性衡量指标,发现窄域微调模型得分差,存在身份混淆等问题,其病态或限制对一致未对齐行为的研究。
AI 中文摘要
大量研究基于输出方差衡量模型一致性,但未充分考虑竞争原因。我们识别出歧义与无差别这两个此类原因,并引入包含175个问题的数据集,其中矛盾答案难以用上述任一原因解释。随后我们通过重采样同一问题的答案时出现的矛盾来衡量不一致性。与其他方法相比,我们的指标具有高特异性,仅在问题明显时才将模型评为不一致。即便如此,我们发现窄域微调模型得分很差。检查我们方法标记的不一致性,发现文献中的模型存在严重问题,如身份混淆、内省失败和合理化解释。这些发现表明,窄域微调引发的病态可能会限制这些模型能告诉我们的关于一致的、未对齐行为的信息。
英文摘要
A large body of research measures model coherence based on output variance without adequately considering competing causes. We identify two such causes, ambiguity and indifference, and we introduce a set of 175 questions where contradicting answers cannot easily be explained by either. We then measure incoherence in terms of contradictions when resampling answers to the same question. In contrast to other methods our metric has high specificity, and only ranks models as incoherent when the issues are glaring. Even so, we find narrow finetunes score poorly. Inspecting inconsistencies flagged by our method, we find that model organisms from the literature display severe issues such as identity conflation, introspection failures and rationalizations. These findings suggest that the pathologies induced by narrow finetuning may limit what these models can tell us about coherent misaligned behaviour.
Comments18 pages, 5 figures, 4 tables. Code and dataset: https://github.com/themachinefan/twominds