arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Codebook Agent:面向大语言模型多智能体系统的摊销式拓扑设计

Codebook Agent: Amortized Topology Design for LLM Multi-Agent Systems

Jinxi Yu, Yubei Li, Eric Hanchen Jiang, Zhi Zhang, Dong Liu, Wenxiao Zhao, Levina Li, Kai-Wei Chang, Ying Nian Wu

arXiv 2609.02264首次发表:更新:

发表机构

University of California, Los Angeles(加利福尼亚大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文针对LLM多智能体系统拓扑设计的现有建模缺陷,提出Codebook Agent方法,该方法通过向量量化自编码器与奖励加权MLP实现高效拓扑生成,在6个基准上准确率最优,推理速度快且令牌消耗少。

AI 中文摘要

针对每个查询自适应调整大语言模型(LLM)多智能体系统的通信拓扑,可同时提升准确率与效率,但当前研究者将该问题视为条件图生成任务:变分、自回归或扩散解码器在N×N邻接矩阵空间中搜索,基于效用与边数等结构损失训练的图网络代理对采样候选拓扑进行排序。本文认为该问题建模方式与实际需求不匹配。实证结果显示,即便码本容量从8增长到64,通过奖励筛选的拓扑仍会坍缩为约6种不同图;边数与测得的令牌消耗呈负相关(皮尔逊相关系数r≈-0.4),因此图稀疏会使推理成本更高;当智能体共享配置时,基于智能体轮廓节点的消息传递评分器具有邻接不变性——这是已发布基准的默认配置,导致其在该场景下完全无法对候选拓扑排序。基于这三个事实,本文提出Codebook Agent:向量量化自编码器将成功的拓扑压缩为与查询无关的16项码本;奖励加权多层感知机(MLP)将查询嵌入映射到码本的分布;读取扁平化邻接矩阵的MLP代理,通过回归测得的效用与每任务归一化令牌成本,在单次批量前向传播中对解码出的顶级候选拓扑重新排序。在测试阶段无需迭代搜索与消息传递,Codebook Agent在所有6个对比基准上均为最准确方法(平均准确率84.6,最强现有设计方法为83.0),生成拓扑耗时2.4毫秒,且LLM令牌使用量减少21.9%至33.2%。

英文摘要

Adapting the communication topology of an LLM multi-agent system to each query improves both accuracy and efficiency, yet current designers treat this as conditional graph generation: a variational, autoregressive, or diffusion decoder searches the $N \times N$ adjacency space, and a graph-network proxy trained on utility and a structural cost such as edge count ranks the sampled candidates. We argue that this formulation is misaligned with the problem. Empirically, topologies that survive a reward filter collapse to about six distinct graphs even when the codebook capacity grows from 8 to 64; edge count is negatively correlated with measured token consumption (Pearson $r \approx -0.4$), so sparsifying the graph makes inference more expensive; and a message-passing scorer over agent-profile nodes is adjacency-invariant whenever agents share a profile---the default configuration of published benchmarks---so it cannot rank candidates at all in that regime. These three facts motivate Codebook Agent: a vector-quantized autoencoder compresses successful topologies into a query-independent 16-entry codebook; a reward-weighted MLP maps the query embedding to a distribution over codes; and an MLP proxy that reads the flattened adjacency, regressed on measured utility and per-task normalized token cost, reranks the top decoded candidates in a single batched forward pass. With no iterative search and no message passing at test time, Codebook Agent is the most accurate method on all six benchmarks we compare (84.6 average against 83.0 for the strongest prior designer), emits a topology in 2.4 ms, and uses 21.9--33.2% fewer LLM tokens.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑