arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

自涌现智能体架构:行为惯性隐马尔可夫模型、反思性元认知与社会对比自我建模

Self-Emergence Agent Architecture:Behavior-Inertia HMM, Reflexive Metacognition,and Social-Contrastive Self-Modeling

Xiaoyang Liu

arXiv 2609.17331首次发表:更新:

AI 中文总结

针对LLM智能体的人格漂移、非进化反思和缺乏自我-他者边界问题,提出SEAA架构,通过HMM编码惯性、反思循环更新参数及多智能体社会比较形成闭环,实现自发对称破缺,涌现出差异化人格与社会结构。

AI 中文摘要

大型语言模型(LLM)智能体展现出强大的语言生成和问题解决能力,但存在三个结构性局限:人格漂移、非进化性反思以及缺乏自我-他者边界。现有的生成式智能体模拟依赖静态记忆和固定提示,既无法维持行为惯性,也无法实现内生性自我进化。我们提出了自涌现智能体架构(SEAA),该架构整合了三个组件:(i)一个隐马尔可夫模型(HMM),将长期行为和认知惯性编码为可编辑的状态转移矩阵;(ii)一个反思式(Reflexion风格)言语元认知循环,其输出直接更新HMM参数本身,而非仅仅作为文本存储;(iii)一个多智能体社会环境,其中初始相同的智能体持续将自身行为与他人行为进行比较。这三个组件形成一个闭环:社会行动→反馈→自我反思→惯性更新→差异化行动。我们提出了三个可证伪的假设,并提供了一个具有操作性指标的可复现实验协议。一个不依赖语言模型的原型表明,该循环会自发地打破对称性:初始相同的智能体巩固出不同且稳定的人格,而匹配的对照组则没有。使用托管LLM进行的实验将这些差异表现为不同的第一人称自我叙事,并且一个五智能体审议过程自发形成了社会结构——一个共识中心和一致被拒绝的离群者——而在对照组中则不存在。遵循受庄子启发的认识论不可知论立场,SEAA仅研究可观察的行为涌现,不对主观感受质作任何断言。本工作贡献了一个统一框架、一个带有伪代码的具体架构、机制性证据,以及一个用于研究人工自我涌现的显微镜式沙盒。

英文摘要

Large language model (LLM) agents exhibit strong language-generation and problem-solving capabilities, yet suffer from three structural limitations: personality drift, non-evolutionary reflection, and the absence of a self-other boundary. Existing generative-agent simulations rely on static memory and fixed prompts, maintaining neither behavioral inertia nor endogenous self-evolution. We propose the Self-Emergence Agent Architecture (SEAA), which integrates three components: (i) a Hidden Markov Model (HMM) that encodes long-term behavioral and cognitive inertia as an editable state-transition matrix; (ii) a Reflexion-style verbal metacognition loop whose output updates the HMM parameters themselves, rather than merely being stored as text; and (iii) a multi-agent social environment in which initially identical agents continuously compare their behavior with others'. The three components form a closed loop: social action $\to$ feedback $\to$ self-reflection $\to$ inertia update $\to$ differentiated action. We state three falsifiable hypotheses and provide a reproducible experimental protocol with operational metrics. A language-model-free prototype shows the loop spontaneously breaks symmetry: initially identical agents consolidate distinct, stable personalities whereas matched controls do not. Experiments with a hosted LLM surface these differences as distinct first-person self-narratives, and a five-agent deliberation spontaneously develops social structure---a consensus hub and a unanimously rejected outlier---absent in the control. Following an epistemologically agnostic stance inspired by Zhuangzi, SEAA studies only observable behavioral emergence and makes no claim about subjective qualia. This work contributes a unified framework, a concrete architecture with pseudocode, mechanistic evidence, and a microscope-style sandbox for studying artificial-self emergence.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑