arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SimSkill:一种用于自主掌握交通模拟的终身学习AI智能体

SimSkill: A Self-Evolving LLM Agent for Skill and Knowledge Accumulation in Traffic Simulation

Qi Liu, Qinzheng Wang, Can Li, Yiming Bie, Wanjing Ma

arXiv 2609.03753首次发表:更新:

发表机构

School of Transportation, Jilin University(吉林大学交通学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出SimSkill智能体,基于SUMO模拟器,通过多类型记忆整合提升交通模拟任务完成率,在三种骨干LLM上提升验证完成率达25个百分点,展示了自然语言与可执行工具结合的设计范式。

AI 中文摘要

随着大语言模型(LLM)的能力日益增强,AI系统的长期价值不仅取决于解决单个请求,还取决于将经验和积累的知识转化为持久、可复用的能力。我们提出SimSkill,这是一种围绕城市移动模拟(SUMO)交通模拟器构建的自进化智能体。SimSkill会识别能力缺口,生成并解决基于环境的任务,通过行动-评论者循环验证解决方案,并将经验整合到情景记忆、程序记忆和语义记忆中,且无需更新骨干模型。通过自主探索,它构建了一个覆盖交通模拟工作流程的可复用库。我们在两个保留的基准上对SimSkill进行评估,使用了三种骨干LLM,并进行了基于人工制品的独立验证。SimSkill将验证后的完成率提高了多达25个百分点,而消融实验显示程序记忆和语义记忆具有互补贡献。其优势取决于骨干模型和计算预算:记忆并非对所有模型都有提升,也未均匀降低推理成本。更广泛地说,SimSkill展示了一种设计范式,其中自然语言保留并组合计算能力,而可执行工具和代码提供精确且可复现的执行。所有代码和实验数据均可在此httpsURL公开获取。

英文摘要

Cumulative culture enables humans to preserve, reuse, and extend knowledge and skills across experiences and generations. Inspired by this principle, we introduce \textit{SimSkill}, a self-evolving agent built around the Simulation of Urban MObility (SUMO) traffic simulator. SimSkill continually identifies capability gaps, generates and solves environment-grounded tasks, verifies solutions through an action--critic loop, and consolidates experience into episodic, procedural, and semantic memory. Through autonomous exploration, it builds a library of reusable skills and knowledge spanning major stages of the traffic-simulation workflow. We evaluate SimSkill on two held-out benchmarks across three backbone LLMs, with each result independently verified. It improves verified success by up to 25 percentage points, and ablations show complementary contributions from procedural and semantic memory. Its benefits remain backbone- and budget-dependent, as memory does not improve every model or uniformly reduce inference cost. More broadly, SimSkill illustrates a natural-language-centered design paradigm for LLM-based agent systems. Its high-level control logic, operating principles, and accumulated knowledge are expressed in natural language, while an LLM integrates them with executable tools and code to realize precise and reproducible execution. All code and experimental data are publicly available at https://github.com/qiliuchn/SimSkill-V1.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑