发表机构
SINTEF; Singapore Management University; Ozyegin University(辛特夫研究所; 新加坡管理大学; 奥兹耶金大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究系统比较了智能体LLM在五项软件工程任务中的性能,发现多智能体虽提升漏洞检测但能耗和延迟大幅增加,轻量配置更优,为可持续开发工具设计提供指南。
AI 中文摘要
大型语言模型(LLM)在软件工程中的应用日益广泛,包括协调多个智能体的智能体系统,但这些系统带来了更高的计算和环境成本。在本文中,我们对智能体LLM系统在五项软件工程任务上进行了全面的实证研究:代码生成、技术债务识别、代码漏洞检测、日志解析和日志分析。对于每项任务,我们比较了从非智能体单查询基线到多智能体工作流的LLM配置,使用了六个开放权重LLM、两种提示策略和三个硬件平台。我们根据准确性、推理延迟和能耗来评估每种配置。我们的结果揭示了智能体复杂性与能源效率之间的显著权衡:多智能体设计平均消耗的能量是非智能体基线的6.36倍,运行时间是其6.07倍,个别任务-硬件组合的最坏情况减速高达160倍。额外智能体带来的准确性提升有限且因任务而异:多智能体提高了平均漏洞检测准确性,但轻量级非智能体和单智能体配置仍然主导帕累托前沿,占66个帕累托最优配置中的59个。模型和提示选择作为任务特定的杠杆,其有效方向因任务而异,而非作为全局默认值。我们将这些发现转化为可持续、任务感知的基于LLM的开发工具的设计指南。
英文摘要
Large language models (LLMs) are increasingly used in software engineering, including agentic systems that coordinate multiple agents, but impose higher computational and environmental costs. In this paper, we present a comprehensive empirical study of agentic LLM systems across five software engineering tasks: code generation, technical debt identification, code vulnerability detection, log parsing, and log analysis. For each task, we compare LLM configurations that range from a non-agentic single-query baseline to multi-agent workflows, using six open-weight LLMs, two prompt strategies, and three hardware platforms. We assess each configuration in terms of accuracy, inference latency, and energy consumption. Our results reveal substantial trade-offs between agentic complexity and energy efficiency: multi-agent designs consume on average 6.36$\times$ as much energy and run 6.07$\times$ as long as the non-agentic baseline, with worst-case slowdowns of up to 160$\times$ for individual task--hardware pairs. Accuracy gains from additional agents are limited and task-specific: multi-agent improves average vulnerability-detection accuracy, but lightweight non-agentic and single-agent configurations still dominate the Pareto front, accounting for 59 of 66 Pareto-optimal configurations. Model and prompt choice act as task-specific levers whose effective direction varies between tasks rather than as global defaults. We translate these findings into design guidelines for sustainable, task-aware LLM-based development tools.
Comments50 pages, 14 figures