发表机构
Télécom SudParis, Institut Polytechnique de Paris; Cardiff Metropolitan University; Aberystwyth University(巴黎电信学院,巴黎综合理工学院; 卡迪夫城市大学; 阿伯里斯特威斯大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究解耦多智能体LLM的拓扑与多样性,在固定预算下评估六种配置,发现并行学习专业化在低资源情感检测中表现最佳,且智能体区分方式比拓扑影响更大。
AI 中文摘要
多智能体大语言模型系统结合了多次推理调用,但先前的工作常常混淆了调用之间的连接方式与调用多样化的方式。我们独立研究这些因素:推理拓扑和智能体间多样性的来源。在一个受控的 $2 \ imes 3$ 矩阵中,我们将并行聚合和顺序细化与随机采样、角色提示和学习的 QLoRA 专业化进行交叉组合,在固定三次调用预算和每个骨干模型内的输出协议下进行。使用 Qwen2.5-14B-Instruct 和 Llama-3.1-8B-Instruct,我们在九种语言的多语言低资源情感检测上评估所有六种配置。并行学习专业化在 Qwen 上最强,达到 52.83 Macro-F1,在 Llama 上达到 52.94。在 Qwen 上,它也超过了同骨干的零样本、少样本、思维链和七次调用自一致性基线。优选的拓扑取决于多样性来源:顺序细化有助于随机和提示设置,而学习到的宽度优势从 Qwen 上的 2.83 分缩小到 Llama 上的 0.17 分。深度分析表明,后期的学习专家可能覆盖早期正确的预测,尽管总体效果依赖于骨干模型。总体而言,智能体的区分方式比拓扑产生更大的性能变化,应结合专业化进行联合评估。
英文摘要
Multi-agent LLM systems combine multiple inference calls, but prior work often confounds how calls are connected with how they are diversified. We study these factors independently: inference topology and source of inter-agent diversity. In a controlled $2 \times 3$ matrix, we cross parallel aggregation and sequential refinement with stochastic sampling, role prompting, and learned QLoRA specialization, under a fixed three-call budget and output protocol within each backbone. Using Qwen2.5-14B-Instruct and Llama-3.1-8B-Instruct, we evaluate all six configurations on multilingual low-resource emotion detection across nine languages. Parallel learned specialization is strongest on Qwen at 52.83 Macro-F1 and reaches 52.94 on Llama. On Qwen it also exceeds same-backbone zero-shot, few-shot, CoT, and seven-call self-consistency baselines. The preferred topology depends on diversity source: sequential refinement helps stochastic and prompted settings, while the learned Width advantage shrinks from 2.83 points on Qwen to 0.17 on Llama. Depth-wise analysis suggests that later learned specialists can overwrite correct early predictions, although the aggregate effect is backbone-dependent. Overall, how agents are differentiated produces larger performance shifts than topology, which should be evaluated jointly with specialization.
Comments23 pages, 5 figures, 25 tables. Accepted at the REALM Workshop at EMNLP 2026. Code: https://github.com/eracoding/topologyxdiversity