Do More Agents Help? Controlled and Protocol-Aligned Evaluation of LLM Agent Workflows
更多智能体有帮助吗?LLM智能体工作流的受控与协议对齐评估
机构 * Beijing University of Posts and Telecommunications(北京邮电大学) ; Westlake University(西湖大学) ; Zhejiang University(浙江大学) ; Duke Kunshan University(杜克大学昆山分校) ; Hong Kong University of Science and Technology(香港科技大学) ; Zhejiang University of Technology(浙江工业大学)
专题命中 推理评测 :reasoning(abstract);分类 cs.AI
AI总结 提出BenchAgent框架,在统一协议下比较单智能体、固定多智能体和演化多智能体工作流,发现大多数多智能体系统在准确率上未超越单智能体基线,但运行时生成的工作流在GAIA上表现优异。
Comments https://github.com/LINs-lab/MASArena/tree/BenchAgent