arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05279cs.AIcs.MA

测试大语言模型(LLM)智能体团队的可互换性

Testing Interchangeability in LLM Agent Teams

Jianxin Gao, Tianyi Yu, Linna Deng, Runze Li, Zining Wang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究验证生产多智能体系统中智能体可互换的假设,通过交换LLM智能体团队角色匹配成员,发现任务得分受影响小但通信量增加,且协调效率的可替代性低于任务结果。

中文摘要 AI 辅助

生产级多智能体系统会持续替换智能体,其前提是,担任某一角色的智能体与任何能胜任该工作的其他智能体是可互换的。我们对这一假设进行了验证。每个设置下独立组建8支团队,均基于同一基础模型完成相同任务,每个智能体在10次组建周期中保留私人笔记本;随后我们在团队间交换角色匹配的智能体,并在保留任务上测量产生的变化。与安慰剂(在不更换人员的情况下重现人员名单变更的干扰)相比,交换对任务得分的影响很小,但使团队每单位进展的通信量增加了16%至63%;在Hanabi游戏中,交换后的智能体比无经验的智能体成本更高,这与其从原合作伙伴处习得的惯例产生干扰一致。在Collab-Overcooked中,当设定议程的智能体被替换时,大部分额外通信来自未被替换的智能体。对基础模型、解码温度和组建周期长度进行的三项消融实验显示,交换代价与另一数量(独立组建的团队之间的差异程度)同步变化:贪心解码会降低两者,将团队的历史记录翻倍会提高两者。在这些设置中,智能体在任务结果方面比在协调效率方面更具可替代性,且在更长的组建历史后交换效应更大。

英文摘要

Production multi-agent systems replace agents constantly, on the assumption that an agent filling a role is interchangeable with any other agent that can do the job. We test that assumption. Eight teams per setting are formed independently from one base model on the same tasks, each agent keeping a private notebook across ten formation episodes; we then trade role-matched agents between teams and measure what changes on held-out tasks. Against a placebo that reproduces the disruption of a roster change without changing who occupies the seat, a swap costs little in task score but raises the communication a team spends per unit of progress by 16 to 63 percent, and in Hanabi a swapped agent is more expensive than an inexperienced one, consistent with interference from conventions learned with its former partner. In Collab-Overcooked, when the agent that sets the agenda is replaced, most of the extra communication comes from the agent that stayed. Three ablations, over base models, decoding temperature and formation length, move the swap penalty alongside one other quantity: how far independently formed teams drift apart. Greedy decoding lowers both; doubling a team's history raises both. In these settings, agents are more fungible in task outcome than in coordination efficiency, with larger swap effects after longer formation histories.

发表机构

  • China Agricultural University(中国农业大学)
  • Tianjin University of Finance and Economics(天津财经大学)
  • Jilin University(吉林大学)
  • Tianjin University of Science and Technology(天津科技大学)

机构由 AI 辅助整理,请以论文原文为准。

↑