多智能体大语言模型系统推理时并行性的两层视角
A Two-Tier Perspective on Inference-Time Parallelism in Multi-Agent LLM Systems
- Beijing University of Posts and Telecommunications(北京邮电大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文提出整合两类并行性的TIPEX框架,在GAIA基准上验证推理时并行性可提升多智能体系统性能,中等难度任务从两类并行性协调中获益最多。
中文摘要 AI 辅助
由大语言模型(LLM)驱动的多智能体系统在推理阶段通常需要多次模型调用和复杂的协调过程,其执行策略直接影响系统的准确率、延迟和计算成本。并行执行是提升推理效率的一种手段。从推理阶段执行的视角出发,本文将多智能体系统中的并行性建模为两类不同层级的决策过程:副本并行性(Replica Parallelism),在任务层面探索多条完整的解决方案路径;结构并行性(Structural Parallelism),通过任务分解在单条解决方案路径内实现并发执行。然而,不同形式并行性的作用及其相互关系,在统一组织与协调方面仍缺乏系统研究。为此,本文提出TIPEX,一种可控执行框架,该框架在统一执行语义下整合上述两类并行性,协调其在推理过程中的作用,同时支持对不同并行策略和参数配置进行系统的组合与分析。在GAIA基准上开展的系统实验表明,推理时并行性可在提升准确率、降低端到端延迟的同时,伴随token消耗的增加。进一步分析显示,副本并行性与结构并行性在不同任务复杂度下呈现互补效应,中等难度任务从二者的协调中获益最多,而过于激进的并行策略未必能带来更好的性能。
英文摘要
Large language model (LLM)-driven multi-agent systems typically require multiple model invocations and complex coordination during inference, and their execution strategies directly affect system accuracy, latency, and computational cost. Parallel execution provides a means to improve inference-time efficiency. From the perspective of inference-time execution, this paper models parallelism in multi-agent systems as two distinct levels of decision processes: Replica Parallelism, which explores multiple complete solution paths at the task level, and Structural Parallelism, which enables concurrent execution within a single solution path through task decomposition. However, the roles of different forms of parallelism and their interrelationships still lack systematic study in terms of unified organization and coordination. We therefore propose TIPEX, a controllable execution framework that unifies these two levels of parallelism and coordinates their roles within the inference process under a unified execution semantics while supporting systematic combinations and analyses of different parallel strategies and parameter configurations. Systematic experiments on the GAIA benchmark demonstrate that inference-time parallelism can significantly improve accuracy and reduce end-to-end latency at the cost of increased token consumption. Further analysis shows that Replica and Structural Parallelism exhibit complementary effects across task complexities, with tasks of intermediate difficulty benefiting most from their coordination, while overly aggressive parallel strategies do not necessarily yield better performance.