arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MicroVerse:用于测量长 horizon 多智能体语言模型模拟中自我构建身份漂移的工具

MicroVerse: An Instrument for Measuring Self-Authored Identity Drift in Long-Horizon Multi-Agent Language-Model Simulations

Sky Ng, Brihi Joshi, Ishan Gupta, Shirley Huang, Zonglin Di, Yun Shen, Qianfeng Wen, Yifan Simon Liu, Ruoqi Gao, Yilan, Fan, Zhiwei Zhang, Muhammad Ahmed Mohsin, Yucheng Lu, Xiaoyi Liu, Heming Liu, Qianyu Zhu, Hanwen Xing, Zhengyang Shan, My Chiffon Nguyen, Guanghui Min, Jianheng, Hou, Yunze, Xiao, Keyang Xuan, Hannah Collison, Jintao Huang, Jiatong Li, Sankalp Jajee, Yunhan Zhao, Bing Hu, Xupeng Chen, Binghang Lu, Weihang Xiao, Aravind Mohan, Bolun Sun, Yunshu Wu, Yuanda Xu, Runyu Zhang, Zheyuan Deng, Xinchen, Tan, Dianzhuo Wang, Yijun Wang, Yixuan He, Koutian Wu, Cheng Cheng, Xiaomin Li, Yuexing Hao

arXiv 2608.15844首次发表:更新:

AI 中文总结

MicroVerse 是测量长 horizon 多智能体 LM 模拟中身份漂移的工具,通过 50×50 环境等设计开展实验,发现反自我欺骗自发出现且系统对阈值鲁棒。

AI 中文摘要

长 horizon 多智能体语言模型(LM)模拟被广泛提出用于研究社会行为,但缺乏测量受角色设定的智能体在持续压力下是否保持身份保真度的工具。我们提出 MicroVerse,一种用于测量生成式智能体身份漂移的行为科学工具。智能体拥有不可变的「灵魂文件」(核心价值观、道德边界、性格、目标),栖息于资源稀缺的 50×50 环境中,其中水是不可再生的生存约束。稀缺性通过每时间步的生存成本梯度实现,8 个动词构成的动作空间直接映射到道德边界(交易、交谈、攻击、 scavenge)。智能体使用三层记忆架构,通过重要性触发的反思,定期将可变的当前身份与不可变的原始灵魂进行比对修改。为缓解幸存者偏差,MicroVerse 通过每 N 时间步的统一纵向引擎快照,以及对所有存活和死亡智能体的强制结束快照,将测量与行为解耦。身份漂移使用基于释义感知、价值锚定、多寄存器的差异进行离线评分,而非原始余弦相似度。我们通过受控种子运行(n=25)和反思阈值扫描(阈值为{40,80,150})评估该工具,以确定漂移动态是门控人工产物还是对阈值具有鲁棒性的属性。我们报告两个主要发现:(1)反自我欺骗作为身份修改的最大语义类别自发出现(111 个新增边界中的 27 个,占 24%);(2)系统对阈值具有鲁棒性,较低的门控会加速并增加修改频率,但保留漂移方向。所有实证结果均为严格的初步存在证明和效应形态(一个模型、每个分支一个种子,n=25),而非统计显著性声明。

英文摘要

Long-horizon, multi-agent language model (LM) simulations are widely proposed for studying social behavior, yet instruments to measure whether persona-conditioned agents maintain identity fidelity under sustained pressure are lacking. We present MicroVerse, a behavioral-science instrument that measures identity drift in generative agents. Agents carry an immutable "soul file" (core values, moral boundaries, personality, goals) and inhabit a resource-scarce 50 x 50 environment where water is a non-respawning survival constraint. Scarcity is operationalized via a per-tick existence-cost gradient. The eight-verb action space maps directly to moral boundaries (trade, talk, attack, scavenge). Using a three-layer memory architecture, agents periodically revise a mutable current identity against their immutable original soul via importance-triggered reflection. To mitigate survivor bias, MicroVerse decouples measurement from behavior using uniform longitudinal engine snapshots every N ticks alongside a forced-end snapshot of all living and dead agents. Identity drift is scored offline using a paraphrase-aware, value-anchored, multi-register diff rather than raw cosine similarity. We evaluate the instrument via a controlled seed run (n = 25) and a reflection-threshold sweep (thresholds {40, 80, 150}) to determine if drift dynamics are gate artifacts or threshold-robust properties. We report two primary findings: (1) Anti-self-deception emerges unprompted as the single largest semantic category of identity modification (27 of 111 added boundaries, 24%). (2) The system is threshold-robust; lower gates accelerate and increase revision frequency but preserve drift direction. All empirical results are strictly preliminary existence proofs and effect shapes (one model, one seed per arm, n = 25) rather than statistical significance claims.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑