arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24215cs.MA

消费级GPU上的Agentopia:采用8B模型的缩减规模长视距移植方案

Agentopia on a Consumer GPU: A Reduced-Scale Long-Horizon Port with an 8B Model

Luo Huan

首次发表
浏览论文内容

中文总结 AI 辅助

本文在单块消费级GPU上实现并评估了采用8B量化模型的缩减规模Agentopia移植方案,验证了其长视距模拟的可行性并发布了相关资源。

中文摘要 AI 辅助

基于大语言模型(LLM)的多智能体社会模拟已展现出引人注目的成果,但Agentopia此前是使用Qwen3.5-397B-A17B在100个智能体、10个模拟年的规模下进行评估的,消费级硬件上缩减规模部署的表现尚不明确。本文中,我们使用Qwen3-8B-AWQ(一种4-bit量化模型)在单块NVIDIA RTX 5070 Ti(12 GB显存)上实现并评估了缩减规模的Agentopia移植方案。针对该设置,我们引入了三项结构调整:(1)系统管理的分层内存压缩;(2)每个模拟日包含四个活动时段;(3)明确的生理与心理健康状态变量。在三次独立的随机运行中,两次运行完成了52周,第三次完成了50周后达到上下文限制,总计154系统周(770智能体周)。没有智能体死亡,也未记录到基于阈值的健康警告;各运行中包含至少一个NO_RESPONSE字段的活动记录占比为10.15%-10.29%。一项52周的关闭内存运行将L2/L3产物生成与分层内存关联起来;另一项10周的对比实验显示,四个每日时段对应2.72倍的最终记录数量和更低的词汇重复率,但缺失字段率更高。这些对比不支持因果行为主张。我们在一个公开分支中发布了经验证的配置、衍生审计、分析脚本、汇总图表数据及我们的实现变更;原始运行数据和初始角色数据因来源未完全明确而未包含在内。

英文摘要

Large language model (LLM)-based multi-agent social simulation has demonstrated compelling results, but Agentopia was evaluated with 100 agents over 10 simulated years using Qwen3.5-397B-A17B, leaving the behavior of reduced-scale deployments on consumer hardware unclear. In this paper, we implement and evaluate a reduced-scale Agentopia port on a single NVIDIA RTX 5070 Ti(12 GB VRAM) using Qwen3-8B-AWQ, a 4-bit quantized model. We introduce three structural adaptations for this setting: (1) system-managed layered memory compression, (2) four activity blocks per simulated day, and (3) explicit physical- and mental-health state variables. Across three independent stochastic runs, two runs completed 52 weeks and the third completed 50 weeks before reaching the context limit, totaling 154 system-weeks (770 agent-weeks). No agent died,and no threshold-based health warning was logged; activity records containing at least one NO_RESPONSE field occurred at rates of 10.15-10.29% across runs. A 52-week memory-off run tied L2/L3 artifact production to layered memory; a separate 10-week comparison associated four daily time blocks with 2.72 times more finalized records and lower lexical duplication, but a higher missing-field rate. These comparisons do not support causal behavioral claims. We release validated configurations, derived audits, analysis scripts, aggregate figure data, and our implementation changes in a public fork; raw runs and initial persona data are excluded because their redistribution provenance is not fully resolved.

补充信息

↑