两次调用胜过五个智能体:针对本地语言模型的多智能体流水线与自我优化的评估
Two Calls Beat Five Agents: Evaluating Multi-Agent Pipelines Against Self-Refinement for Local Language Models
浏览论文内容
中文总结 AI 辅助
本研究对比多智能体流水线与自我优化方法,发现针对本地7B模型,两次调用自我优化在GSM8K表现更优且令牌用量低,经任务感知重设计后在HumanEval也能维持高准确率,简单方法优于多智能体架构。
中文摘要 AI 辅助
多智能体大语言模型(LLM)流水线系统将任务拆解为多个角色以提升推理能力,但其基准测试主要针对大规模商业模型。本研究将包含五个角色的结构化多智能体系统Parishad部署在本地模型Qwen2.5-7B-Instruct上,在GSM8K(500个问题)和HumanEval(164个问题)两个数据集上,与直接提示及两次调用自我优化方法对比。多智能体系统采用JSON数据格式时,GSM8K准确率从75.0%降至45.0%,归因于误差累积问题;采用明文格式时,准确率恢复至82.0%。两次调用自我优化策略(V1)在GSM8K上可达到86.2%的准确率,令牌使用量降低7.4倍,但在直接准确率已达96.3%的HumanEval上,该V1实现反而使性能降至66.5%;针对HumanEval的任务感知门控重设计(V2)将准确率维持在95.1%。研究结果表明,通信格式与实现细节比架构复杂性更能决定结果,对于本地7B模型部署,更简单的方法可达到或超过多智能体流水线的效果,所有代码与数据均已公开。
英文摘要
Multi-agent LLM pipeline systems break down the task among multiple roles for better reasoning, but are benchmarked mainly with large-scale commercial models. In this study, we investigate Parishad, a structured multi-agent system involving five roles, by deploying it on Qwen2.5-7B-Instruct, a local model, on two datasets: GSM8K (500 questions) and HumanEval (164 questions), compared with prompting directly and two-call self-refinement. The multi-agent system drops GSM8K accuracy from 75.0\% to 45.0\% with JSON data format due to the error accumulation problem. With plaintext format, the accuracy is restored to 82.0\%. A two-call self-refinement strategy (V1) can achieve 86.2\% accuracy on GSM8K, with 7.4$\times$ lower token usage. However, the same V1 implementation on HumanEval---where direct accuracy is already 96.3\%---actively destroys performance (66.5\%). A task-aware gated redesign (V2) applied to HumanEval preserves accuracy at 95.1\%. Our results demonstrate that communication format and implementation details determine outcomes more than architectural complexity, and that simpler approaches match or outperform multi-agent pipelines for local 7B model deployment. All code and data are released.