arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

奖励引导自回归图生成的高效多智能体通信拓扑设计

Reward-Guided Autoregressive Graph Generation for Efficient Multi-Agent Communication Topology Design

Poomphob Suwannapichat, Boonyarit Changaival, Caesar Wu, Pascal Bouvry

arXiv 2608.20099首次发表:更新:

发表机构

University of Luxembourg; King Mongkut’s University of Technology Thonburi(卢森堡大学; 国王蒙kut理工大学thonburi校区)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多智能体系统通信拓扑设计中 token 消耗过高的问题,本文提出 RGA-Designer 方法,通过结合 RLHF 训练奖励模型微调图生成器,在保持任务准确率的同时降低了 token 消耗。

AI 中文摘要

基于大语言模型(LLM)的多智能体系统(MAS)通过协调多个智能体在复杂推理任务上取得了优异性能,但代价是产生了大量的 token 消耗。近期的自动拓扑设计工作 ARG-Designer 将该问题重新表述为自回归图生成任务,然而其训练目标未明确激励模型生成稀疏且高效的拓扑。针对这一局限,本文引入受人类反馈强化学习(RLHF)启发的奖励引导自回归图生成方法 RGA-Designer。我们训练一个奖励模型,该模型同时捕捉任务正确性与结构紧凑性,随后利用该奖励模型作为反馈对预训练的图生成器进行微调。我们的方法在保持与 ARG-Designer 相当的任务准确率的同时,将 token 消耗平均降低了 20.5%。

英文摘要

LLM-based Multi-Agent Systems (MAS) achieve strong performance on complex reasoning tasks by coordinating multiple agents, but at the cost of substantial token consumption. Recent work on automatic topology design, ARG-Designer, has reframed this problem as autoregressive graph generation. However, its training objective provides no explicit incentive for the model to generate sparse and efficient topologies. We address this limitation by introducing a Reward-Guided Autoregressive Graph Generation (RGA-Designer) inspired by Reinforcement Learning from Human Feedback (RLHF). We train a reward model that jointly captures task correctness and structural compactness, and then fine-tune the pretrained graph generator using the reward model as feedback. Our method preserves task accuracy at the level of ARG-Designer while reducing token consumption by an average of 20.5%.

CommentsFull version of extended abstract accepted at ICONIP 2026 (poster)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑