arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

但 AI 智能体将如何运营一个城镇的经济?

But How Would AI Agents Run a Town's Economy?

Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun

arXiv 2609.11108首次发表:更新:

发表机构

Karela Technologies Inc.(卡雷拉科技有限公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过多智能体模拟发现,LLM智能体经济中货币流动在工资和定价环节停滞,财富分配随时间缓慢变化,且模型选择和内存对结果影响不同。

AI 中文摘要

我们将100个配备内存的大型语言模型(LLM)智能体置于一个基于真实博卡拉湖畔地理位置的封闭、货币守恒的空间经济中(负责赚取工资、经营企业、设定价格),并运行该多智能体模拟长达26个模拟周,远超智能体社会研究中典型的1-2周。在91次验证运行中(244万次智能体决策,215亿个令牌),货币以一种特定且可测量的方式停止流动。12倍的旅游需求冲击使企业收入增加4.62倍($p<0.001$),我们将其精确分解为1.50倍的广延边际(更多企业进行交易)和3.07倍的集约边际(每个企业收入增加)。货币传导在此停止。工资变动1.03倍($p=0.42$);3,981个菜单项中仅有0.3%曾被重新定价($p=0.47$)。一次随机现金转移(向100个智能体中的20个发放5,000尼泊尔卢比)从相反方向显示了相同模式:311个脉冲后仍有96.7%被持有,边际消费倾向通过两种独立测量均为3-4%,与零无显著差异。因此,在该文献使用的时间范围内,财富分配几乎冻结(2个模拟周内$\ ho=0.964$),但并非完全冻结。$\ ho$在12周时降至0.832,在26周时降至0.752,这种对时间范围的依赖性在短期研究中无法观察到。匹配的消融实验显示了哪个旋钮真正重要。更换底层LLM会改变我们测量的每个结果($p=0.0039$);删除智能体的内存则不会对任何结果产生可检测的影响。一个纯社交工具在两种模型家族中失败率为94-97%,而经济工具的成功率约为96%,且没有可测量的行为偏移。每个关键数字都经过两次验证:一次由实时验证器进行,另一次由离线重新计算进行,该计算将每个智能体的财富与其自身的签名交易历史进行核对,我们发布完整的运行语料库以供重新分析。

英文摘要

We placed 100 memory-equipped large language model (LLM) agents in charge of a closed, money-conserving spatial economy on real Pokhara Lakeside geography (earning wages, running businesses, setting prices) and ran this multi-agent simulation for up to 26 simulated weeks, well past the 1-2 weeks typical of agent-society studies. Across 91 validated runs (2.44M agent decisions, 21.5B tokens), the money stops moving, in a specific and measurable way. A 12x tourist demand shock raises business revenue 4.62x ($p<0.001$), which we decompose exactly into a 1.50x extensive margin (more businesses trading) and a 3.07x intensive margin (more revenue each). Monetary transmission stops there. Wages move 1.03x ($p=0.42$); 0.3% of 3,981 menu items are ever repriced ($p=0.47$). A randomized cash transfer (NPR 5,000 to 20 of 100 agents) shows the same pattern from the opposite direction: 96.7% is still held 311 pulses later, marginal propensity to consume 3-4% by two independent measures, indistinguishable from zero. The wealth distribution is consequently near-frozen at the horizon this literature uses ($ρ=0.964$ over 2 simulated weeks), but not frozen. $ρ$ falls to 0.832 at 12 weeks and 0.752 at 26, a horizon-dependence no short study can see. Matched ablations show which knob actually matters. Swapping the backing LLM moves every outcome we measure ($p=0.0039$); deleting agents' memory moves none of them detectably. A purely social tool fails 94-97% of the time across two model families, compared with ~96% success on economic tools, with no measurable shift away from it. Every headline number is verified twice, by a live validator and by an offline recomputation that reconciles each agent's wealth against its own signed transaction history, and we release the full run corpus for reanalysis.

Comments8 pages, 7 figures, 6 tables. Dataset and analysis code: https://huggingface.co/datasets/sajalregmi4/agent-town-economy

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑