AI 中文总结
MicroVerse 是测量长 horizon 多智能体 LM 模拟中身份漂移的工具,通过 50×50 环境等设计开展实验,发现反自我欺骗自发出现且系统对阈值鲁棒。
AI 中文摘要
长 horizon 多智能体语言模型(LM)模拟被广泛提出用于研究社会行为,但缺乏测量受角色设定的智能体在持续压力下是否保持身份保真度的工具。我们提出 MicroVerse,一种用于测量生成式智能体身份漂移的行为科学工具。智能体拥有不可变的「灵魂文件」(核心价值观、道德边界、性格、目标),栖息于资源稀缺的 50×50 环境中,其中水是不可再生的生存约束。稀缺性通过每时间步的生存成本梯度实现,8 个动词构成的动作空间直接映射到道德边界(交易、交谈、攻击、 scavenge)。智能体使用三层记忆架构,通过重要性触发的反思,定期将可变的当前身份与不可变的原始灵魂进行比对修改。为缓解幸存者偏差,MicroVerse 通过每 N 时间步的统一纵向引擎快照,以及对所有存活和死亡智能体的强制结束快照,将测量与行为解耦。身份漂移使用基于释义感知、价值锚定、多寄存器的差异进行离线评分,而非原始余弦相似度。我们通过受控种子运行(n=25)和反思阈值扫描(阈值为{40,80,150})评估该工具,以确定漂移动态是门控人工产物还是对阈值具有鲁棒性的属性。我们报告两个主要发现:(1)反自我欺骗作为身份修改的最大语义类别自发出现(111 个新增边界中的 27 个,占 24%);(2)系统对阈值具有鲁棒性,较低的门控会加速并增加修改频率,但保留漂移方向。所有实证结果均为严格的初步存在证明和效应形态(一个模型、每个分支一个种子,n=25),而非统计显著性声明。
英文摘要
Long-horizon, multi-agent language model (LM) simulations are widely proposed for studying social behavior, yet instruments to measure whether persona-conditioned agents maintain identity fidelity under sustained pressure are lacking. We present MicroVerse, a behavioral-science instrument that measures identity drift in generative agents. Agents carry an immutable "soul file" (core values, moral boundaries, personality, goals) and inhabit a resource-scarce 50 x 50 environment where water is a non-respawning survival constraint. Scarcity is operationalized via a per-tick existence-cost gradient. The eight-verb action space maps directly to moral boundaries (trade, talk, attack, scavenge). Using a three-layer memory architecture, agents periodically revise a mutable current identity against their immutable original soul via importance-triggered reflection. To mitigate survivor bias, MicroVerse decouples measurement from behavior using uniform longitudinal engine snapshots every N ticks alongside a forced-end snapshot of all living and dead agents. Identity drift is scored offline using a paraphrase-aware, value-anchored, multi-register diff rather than raw cosine similarity. We evaluate the instrument via a controlled seed run (n = 25) and a reflection-threshold sweep (thresholds {40, 80, 150}) to determine if drift dynamics are gate artifacts or threshold-robust properties. We report two primary findings: (1) Anti-self-deception emerges unprompted as the single largest semantic category of identity modification (27 of 111 added boundaries, 24%). (2) The system is threshold-robust; lower gates accelerate and increase revision frequency but preserve drift direction. All empirical results are strictly preliminary existence proofs and effect shapes (one model, one seed per arm, n = 25) rather than statistical significance claims.