AI 中文总结
该研究提出用低成本拟合的低参数模型替代LLM智能体,可在笔记本电脑上模拟任意数量的LLM智能体群体,通过「交互顺序×记忆」分类法验证了该方法在多个LLM模拟上的有效性。
AI 中文摘要
模拟大量大语言模型(LLM)智能体的群体成本高昂,但这类模拟所关注的问题通常是宏观层面的:如相行为、典型事实以及随智能体数量N的缩放情况,而非单个智能体的认知情况。我们将统计物理学的观察转化为一种方法:用一个由数百到数千次低成本查询拟合出的低参数模型替代每个LLM智能体,随后可在笔记本电脑上运行任意数量N的智能体群体。该方法是否有效主要取决于每个智能体的感知情况,且可在模拟运行前判定。我们引入了「交互顺序×记忆」分类法,将感知与记忆映射到有效理论,并预测代理误差的N缩放趋势。我们在对LLM宏观经济模型EconAgent的忠实复现以及另外7个已命名的LLM模拟上验证了该分类法,其中智能体决策来自真实LLM的 elicitations(主要为DeepSeek),成本仅几美元;预测的误差趋势在每个单元上均成立,且两项被反驳的预测(均针对强饱和响应,可追溯至其曲率)也被无自由参数的理论定量匹配。
英文摘要
Simulating societies of many large language model (LLM) agents is expensive, yet the questions asked of such simulations are usually macroscopic: phase behaviour, stylised facts, and scaling with the number of agents $N$, not the cognition of any single agent. We turn a statistical-physics observation into a method: replace each LLM agent by a low-parameter model fitted from a few hundred to a few thousand cheap queries, then run the society at any $N$ on a laptop. Whether this works is decided before the simulation runs, chiefly by what each agent perceives. We introduce an [interaction order x memory] taxonomy that maps perception and memory to an effective theory and a predicted $N$-trend of the surrogate error. We validate it on a faithful reimplementation of the LLM macroeconomy EconAgent and seven further named LLM simulations, with agent decisions cloned from genuine LLM elicitations (primarily DeepSeek) for a few dollars; the predicted error trends hold cell by cell, and the two refuted predictions, both on a strongly saturating response and traced to its curvature, are themselves matched quantitatively by the theory with no free parameters.
Comments25 pages, 12 figures. Code and data at github.com/YehudaItkin/poor-mans-agentic-modeling; systematic review and pre-registration archived at Zenodo (doi:10.5281/zenodo.21198322, doi:10.5281/zenodo.21340310)