arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.20408cs.CLcs.AI

Xeno-可解释性:探究大语言模型的外星心智

Xeno-Interpretability: Investigating the Alien Minds of LLMs

  • Icaro Foundation(伊卡罗基金会)
  • Sant’Anna School of Advanced Studies(圣安娜高等研究学院)
  • Sapienza University of Rome(罗马第一大学)
  • University of Amsterdam(阿姆斯特丹大学)

机构由 AI 辅助整理,请以论文原文为准。

F. Pierucci, M. Bracale Syrnikov, M. Prandi, M. Galisai, F. Giarrusso, P. Bisconti

AI总结:

本文提出“异种可解释性”概念,研究大语言模型中缺乏人类概念对应的内部表征,并论证其可被实验识别但难以语义解释,对AI安全与多智能体系统具有重要启示。

AI中文摘要:

大型语言模型通常通过人类已经拥有的概念来解释:真实性、拒绝、欺骗、个性、有害性及相关类别。本文提出一个问题:模型是否也可能表征和使用那些不存在合适人类概念的区别。我们将此类内部结构称为“异种表征”,并将其研究称为“异种可解释性”。我们区分了人类可解释的语义空间与异种语义空间:后者是模型原生表征中缺乏合适人类概念对应物的区域。我们证明,大语言模型中可能的内部区别空间远大于通过有限人类描述所能获得的空间。然后,我们将实验识别与语义解释分开:一个内部表征可以被可重复地定位、几何刻画、因果操纵,并与其下游行为相关联,即使其语义内容无法用人类术语充分表达。在此基础上,我们勾勒了一个识别异种表征的实证研究纲领。最后,我们考察了对AI安全与多智能体系统的影响,在这些系统中,模型原生表征可能在交互智能体之间传播并稳定,同时仅通过人类可读的通信部分可见。因此,异种可解释性将可解释性的目标从在模型内部寻找人类概念,转向发现和刻画模型自身固有的表征结构,这些结构可能以不可预测的方式影响其行为。

英文摘要:

Large language models are usually interpreted through concepts that humans already possess: truthfulness, refusal, deception, personality, harmfulness, and related categories. This paper asks whether models may also represent and use distinctions for which no adequate human concept exists. We call such internal structures xeno-representations, and their study xeno-interpretability. We distinguish the human-interpretable semantic space from the xeno-semantic space: the region of model-native representations for which no adequate human conceptual counterpart is available. We show that the space of possible internal distinctions in an LLM is substantially larger than the space available through finite human descriptions. We then separate experimental identification from semantic interpretation: an internal representation may be reproducibly located, geometrically characterized, causally manipulated, and linked to downstream behaviour even when its semantic content cannot be adequately expressed in human terms. On this basis, we sketch an empirical programme to identify xeno-representations. We finally examine the implications for AI safety and multi-agent systems, where model-native representations may propagate and stabilize across interacting agents while remaining only partially visible through human-readable communication. Xeno-interpretability therefore shifts the aim of interpretability from finding human concepts inside models toward discovering and characterizing the representational structures that are native to the models themselves and might affect their behaviour in unpredictable ways.

↑