arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将小型大语言模型(LLM)训练为空间多智能体策略

Training Small LLMs as Spatial Multi-Agent Policies

Yi Mao, Andrew Perrault

arXiv 2608.01425首次发表:更新:

发表机构

The Ohio State University(俄亥俄州立大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对小型冻结LLM在空间合作博弈中表现不佳的问题,通过机械合成可行性守卫构建符号选项库,用PA-MAGRPO训练LoRA适配器,实现了小型LLM的有效多智能体策略,并揭示奖励与合作解耦,需结合行为评估。

AI 中文摘要

基于大语言模型(LLM)的多智能体系统与多智能体强化学习的训练正迅速受到关注,同时有并行研究指出,这类系统的评判应基于其行为而非仅基于奖励。我们在空间合作博弈中同时开展这两方面的研究:小型冻结LLM经低级动作提示后完全失败,获得零奖励。我们基于选项/半马尔可夫决策过程(semi-MDP)框架,以及因智能体间选项执行异步性而衍生的宏动作分散式部分可观测马尔可夫决策过程(macro-action Dec-POMDP)多智能体扩展框架,为每个博弈配备符号选项库:由符号规划器执行的带类型、状态可行、短视距的行为。每个库由前沿编码模型从博弈源代码生成;过滤每个选项菜单的可行性守卫则通过低成本随机策略预演回滚机械合成:仅当守卫能解释重复执行失败且不遗漏记录的成功时才被采用,因此无需手动编写、选择或对守卫进行奖励调优。每个智能体的LLM作为其选项策略,配备由多智能体GRPO(PA-MAGRPO)的智能体专属变体训练的智能体专属LoRA适配器;这使冻结基础模型在三款博弈和四款小型骨干模型上从零奖励提升至具备竞争力的表现。行为审计进一步揭示,奖励与合作存在解耦:奖励曲线上升可能仅意味着一个智能体学会独自完成整个任务,而其伙伴闲置——合作仅在任务需要时才会出现。因此,仅靠奖励无法可靠反映合作情况,行为评估必须与奖励评估并行开展。

英文摘要

Training LLM-based multi-agent systems with multi-agent reinforcement learning is rapidly gaining traction, and a parallel line of work argues that such systems should be judged by their behavior, not only their reward. We take up both threads in spatial cooperative games, where small frozen LLMs prompted with low-level actions fail outright, earning zero reward. Guided by the options/semi-MDP framework---and, because option execution is asynchronous across agents, its multi-agent extension in macro-action Dec-POMDPs---we equip each game with a library of symbolic \emph{options}: typed, state-feasible, short-horizon behaviors executed by a symbolic planner. Each library is drafted by a frontier coding model from the game's source code; the feasibility guards that filter each menu are then synthesized mechanically from cheap random-policy burn-in rollouts---a guard is adopted only if it explains repeated execution failures while hiding no logged success---so no guard is authored, selected, or reward-tuned by hand. Each agent's LLM acts as its policy over options, with a private per-agent LoRA adapter trained by a per-agent variant of multi-agent GRPO (PA-MAGRPO); this lifts frozen bases from zero reward to competent play across three games and four small backbones. Behavioral audits then reveal that reward and cooperation decouple: a rising reward curve may simply mean that one agent has learned to run the entire task alone while its partner idles---cooperation emerges only when the task makes it necessary. Reward alone is thus an unreliable readout of cooperation; behavioral evaluation must sit alongside it.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑