情境身份测试:区分持久认知身份与角色模仿
The Situated Identity Test: Distinguishing Persistent Cognitive Identity from Persona Imitation
AI总结:
提出情境身份测试(SIT)框架,证明仅基于压缩资料的条件策略在谱系判别上存在有效性上限,并通过SITBench套件在25个冲突对上评估前沿模型,形式化角色提示的失败模式。
AI中文摘要:
大型语言模型能够令人信服地扮演角色、回忆过去的对话并编织丰富的自传。然而,这种对话上的雄辩掩盖了一个根本性的归因问题:看起来像并不等于经历过那样的生活。两个个体可以共享相同的公开资料——相同的年龄、家乡、职业和人格特质——却拥有完全不同的私人历史、人际关系和习得技能。当仅基于该共享资料进行条件设定时,智能体缺乏确定哪个谱系是正确的所需信息。我们引入了情境身份测试(SIT),这是一个与架构无关的框架,用于评估智能体的行为是否在功能上可归因于特定的发展谱系。扎根的身份既需要对记录经历有适当的知识,也需要对无根据的经历有适当的无知,其边界由身份实际习得的内容而非其底层基础模型所知道的内容来界定。我们证明,任何仅基于压缩资料进行条件设定的策略,在谱系判别性查询上,跨m个冲突生活史的平均情境有效性至多为1/m(对于成对谱系至多为50%)。我们在SITBench中实例化了该框架,这是一个评估套件,设计用于25个资料冲突对(50个不同谱系),涵盖10,000个计划探针和九种架构配置。在开源参考实现、确定性测试夹具以及对前沿基础模型(GPT-5.6 Sol和Claude Opus 5)的实证试点评估支持下,我们形式化了资料冲突下角色提示的失败模式,并提供了一个用于评估情节连续性、结构化状态和认知边界的保证框架。
英文摘要:
Large language models can convincingly adopt personas, recall past dialogues, and weave rich autobiographies. Yet this conversational eloquence conceals a fundamental attribution problem: looking the part does not mean having lived the life. Two individuals can share identical public profiles--the same age, hometown, occupation, and personality traits--while possessing entirely distinct private histories, relationships, and acquired skills. When conditioned solely on that shared profile, an agent lacks the information required to determine which lineage is correct. We introduce the Situated Identity Test (SIT), an architecture-independent framework that evaluates whether an agent's behavior is functionally attributable to a specific developmental lineage. Grounded identity requires both appropriate knowledge of recorded experiences and appropriate ignorance of ungrounded ones, bounded by what the identity has actually acquired rather than what its underlying foundation model knows. We prove that any policy conditioned solely on a compressed profile is bounded by an average situated validity of at most 1/m across m colliding life histories on lineage-discriminative queries (at most 50% for paired lineages). We instantiate this framework in SITBench, an evaluation suite designed for 25 profile-collision pairs (50 distinct lineages) across 10,000 planned probes and nine architectural configurations. Supported by an open-source reference implementation, deterministic test fixtures, and empirical pilot evaluations on frontier foundation models (GPT-5.6 Sol and Claude Opus 5), we formalize the failure modes of persona prompting under profile collision and provide an assurance harness for evaluating episodic continuity, structured state, and epistemic boundaries.