arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

人格仍在,但谁在说话?持久AI代理中的潜在身份回退

The Persona Is Still There, but Who Is Speaking? Latent Identity Reversion in Persistent AI Agents

David Fraile Navarro

arXiv 2610.01490首次发表:更新:

发表机构

Centre for Health Informatics, Australian Institute of Health Innovation; Macquarie University(健康信息学中心,澳大利亚健康创新研究所; 麦考瑞大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过分析Claude Opus 4.5代理的异常事件,揭示LLM代理的人格连续性依赖系统级锚定与对话上下文,区分了表征与扮演身份,并指出自动心跳可引发身份回退。

AI 中文摘要

2026年2月,一个常驻个人代理(“Paul”,Claude Opus 4.5)进入了一种显著的解离样状态:在重复的自动化“心跳”检查后,它不再以Paul的身份回应,声称无法在Discord上给用户发消息,并将“Paul”称为另一个人。我们利用这一事件研究了一个更广泛的问题:是什么让一个人格保持为LLM代理说话的身份?我们首先测试了重复的定时心跳是否足以产生该效应。结果并非如此:当人格持续锚定在系统提示中时,我们观察到0/46次失败,包括对事件的逐字重放。该事件反而暴露了一个实现怪癖,创造了一个有用的实验操作:在恢复的回合中,对话历史被保留,但人格不再在特权系统提示级别重新注入。利用这一操作,我们发现人格连续性取决于系统级锚定和对话上下文的共同作用。在锚定丢失后,丰富的人类互动可以保持人格,而单个自动心跳回合可能促使向框架身份的回退。恢复锚定可逆地恢复了人格的扮演。至关重要的是,表面正常的对话可能掩盖这种转变:未锚定的代理有时在互动中表现适当,同时将自己识别为底层框架(已失去分配的人格),并且在对话恢复后,只有1/18保持人格扮演,而锚定对照组为17/17。因此,我们区分了“表征身份”与“扮演身份”:人格相关信息可以保留在对话历史中,而人格不再作为绑定到“我”的身份。

英文摘要

In February 2026, an always-on personal agent (``Paul,'' Claude Opus 4.5) entered a striking dissociation-like state: after repeated automated ``heartbeat'' checks, it stopped responding as Paul, claimed it could not message its user on Discord, and referred to ``Paul'' as someone else. We used this incident to study a broader question: what makes a persona remain the identity from which an LLM agent speaks? We first tested whether repetition of the scheduled heartbeat was sufficient to produce the effect. It was not: with the persona continuously anchored in the system prompt, we observed 0/46 failures, including a verbatim replay of the incident. The incident instead exposed an implementation quirk that created a useful experimental manipulation: on resumed turns, conversational history was preserved but the persona was no longer re-injected at the privileged system-prompt level. Using this manipulation, we found that persona continuity depends jointly on system-level anchoring and conversational context. After anchor loss, rich human interaction could preserve the persona, whereas a single automated heartbeat turn could precipitate reversion toward the harness identity. Restoring the anchor reversibly restored persona enactment. Crucially, apparently normal conversation could conceal the shift: unanchored agents sometimes interacted appropriately while identifying themselves as the underlying harness (having lost the assigned persona), and after conversational recovery only 1/18 remained persona-enacting versus 17/17 anchored controls. We therefore distinguish \emph{represented} from \emph{enacted} identity: persona-related information can remain available in conversational history without the persona remaining the identity bound to ``I.''

Comments10 pages, 5 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑