arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

工程可持续智能体:面向开发者工作流的智能体LLM系统比较研究

Engineering Sustainable Agents: A Systematic Comparison of Agentic LLMs for Developer Workflows

Merve Astekin, Yan Naing Tun, Arda Goknil, Erik Johannes Husom, Lwin Khin Shar, Hasan Sözer, Ratnadira Widyasari, Hui Song

arXiv 2610.03010首次发表:更新:

发表机构

SINTEF; Singapore Management University; Ozyegin University(辛特夫研究所; 新加坡管理大学; 奥兹耶金大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究系统比较了智能体LLM在五项软件工程任务中的性能,发现多智能体虽提升漏洞检测但能耗和延迟大幅增加,轻量配置更优,为可持续开发工具设计提供指南。

AI 中文摘要

大型语言模型(LLM)在软件工程中的应用日益广泛,包括协调多个智能体的智能体系统,但这些系统带来了更高的计算和环境成本。在本文中,我们对智能体LLM系统在五项软件工程任务上进行了全面的实证研究:代码生成、技术债务识别、代码漏洞检测、日志解析和日志分析。对于每项任务,我们比较了从非智能体单查询基线到多智能体工作流的LLM配置,使用了六个开放权重LLM、两种提示策略和三个硬件平台。我们根据准确性、推理延迟和能耗来评估每种配置。我们的结果揭示了智能体复杂性与能源效率之间的显著权衡:多智能体设计平均消耗的能量是非智能体基线的6.36倍,运行时间是其6.07倍,个别任务-硬件组合的最坏情况减速高达160倍。额外智能体带来的准确性提升有限且因任务而异:多智能体提高了平均漏洞检测准确性,但轻量级非智能体和单智能体配置仍然主导帕累托前沿,占66个帕累托最优配置中的59个。模型和提示选择作为任务特定的杠杆,其有效方向因任务而异,而非作为全局默认值。我们将这些发现转化为可持续、任务感知的基于LLM的开发工具的设计指南。

英文摘要

Large language models (LLMs) are increasingly used in software engineering, including agentic systems that coordinate multiple agents, but impose higher computational and environmental costs. In this paper, we present a comprehensive empirical study of agentic LLM systems across five software engineering tasks: code generation, technical debt identification, code vulnerability detection, log parsing, and log analysis. For each task, we compare LLM configurations that range from a non-agentic single-query baseline to multi-agent workflows, using six open-weight LLMs, two prompt strategies, and three hardware platforms. We assess each configuration in terms of accuracy, inference latency, and energy consumption. Our results reveal substantial trade-offs between agentic complexity and energy efficiency: multi-agent designs consume on average 6.36$\times$ as much energy and run 6.07$\times$ as long as the non-agentic baseline, with worst-case slowdowns of up to 160$\times$ for individual task--hardware pairs. Accuracy gains from additional agents are limited and task-specific: multi-agent improves average vulnerability-detection accuracy, but lightweight non-agentic and single-agent configurations still dominate the Pareto front, accounting for 59 of 66 Pareto-optimal configurations. Model and prompt choice act as task-specific levers whose effective direction varies between tasks rather than as global defaults. We translate these findings into design guidelines for sustainable, task-aware LLM-based development tools.

Comments50 pages, 14 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑