arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

迈向基于智能体的智能体模型:可行性、性能和统计模型检查

Towards Agentic Agent-based Models: Feasibility, Performance, and Statistical Model Checking

Stefano Blando, Emanuele Guerrazzi, Riccardo Porcedda, Giuseppe Squillace, Max Tschaikowski, Andrea Vandin

arXiv 2607.17948首次发表:更新:

发表机构

Sant’Anna School of Advanced Studies Pisa; Sapienza University of Rome; DTU Technical University of Denmark; University of Pisa(比萨圣安娜高等研究学校; 罗马第一大学; 丹麦技术大学; 比萨大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究引入大语言模型驱动决策对基于智能体模型模拟的影响,以Mesa ABM模型为对象,扩展谢林隔离模型,通过统计模型检查进行分析,实验表明较小本地服务LLM有问题,较大模型通过初步检查,还探讨了统计模型检查的作用。

AI 中文摘要

基于智能体的模型(ABM)依靠简单、明确且可重复的规则进行个体决策,复杂的集体行为则源于智能体间的交互。大语言模型(LLM)的进展使得人们想用基于LLM的智能能力来替换、丰富或扰动这些规则。本文研究了这对ABM模拟的可靠性、计算成本和行为的影响。以Mesa ABM模型为研究对象,通过统计模型检查进行分析。扩展了经典的谢林隔离模型,将普通智能体和由LLM驱动的智能体混合。报告了不同大小本地服务LLM的初步实验,较小模型可能在简单语义分类实验中失败或在重复工具调用生成时无法使用,而较大模型通过了初步检查。还讨论了统计模型检查如何估计经典ABM可观测量并量化引入基于LLM的智能体组件对模拟模型的影响。

英文摘要

Agent-based models (ABMs) rely on simple, explicit and reproducible rules for individual decision making, while complex collective behavior emerges from interactions among agents. Recent advances in large language models (LLMs) make it tempting to replace, enrich, or perturb these rules with LLM-based agentic capabilities. However, this raises a methodological question: how does introducing LLM-driven decisions affect the reliability, computational cost, and behavior of ABM simulations? We investigate this for Mesa ABM models, a popular Python library for ABMs, analyzed by statistical model checking. Building on Mesa's integration with the statistical model checker MultiVeStA, we extend the classical Schelling segregation model with a hybrid population: ordinary agents classify neighbors using the standard symbolic rule, while one agent delegates this task to an LLM through tool calls. The LLM-enabled agent receives natural-language descriptions of neighboring agents and invokes tools that increment counters of similar/different neighbors; these counters determine its happiness according to the original Schelling dynamics. This provides a minimal but controlled setting where the semantic, operational, and computational behavior of LLM-based decisions can be studied inside an otherwise standard ABM. We report preliminary experiments with locally served LLMs of different sizes, showing that smaller models may fail simple semantic classification experiments or become operationally unusable during repeated tool-call generation, while larger tested models pass these preliminary checks. We discuss how statistical model checking can estimate classical ABM observables and quantify the impact of introducing agentic LLM components into simulation models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑