arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Agent4RE:用于端到端软件需求工程与基准测试的自优化多智能体框架

Agent4RE: A Self-Refining Multi-agent Framework for End-to-End Software Requirements Engineering and Benchmarking

Yongjian Tang, Linhan Li, Thomas Runkler

arXiv 2610.10628首次发表:更新:

发表机构

Siemens AG; Technical University of Munich(西门子股份公司; 慕尼黑工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出自优化多智能体框架Agent4RE,构建端到端RE基准数据集RE-E2E,其三种变体在8种LLM评估中均优于基线,增强变体评分最高,可作为工业端到端RE实用方案。

AI 中文摘要

现有基于大语言模型(LLM)的软件需求工程(RE)方法通常依赖基础提示策略或初步的智能体协作,未能充分发挥多智能体系统的全部潜力。同时,现有数据集聚焦于孤立的子任务,如需求提取、分类和完整性检测,缺乏从需求启发到生成的端到端RE基准。我们提出Agent4RE——一个自优化多智能体RE系统,它协调专门的智能体并包含两个迭代改进循环。为支持评估,我们构建了RE-E2E——一个基于人类编写的需求规格说明的真实世界数据集,可对RE工作流进行端到端评估。在此基础上,我们进一步提出两个增强版Agent4RE,分别融入自主自优化或结构化人类反馈,并分析它们在不同场景下的优势与局限。对8种大语言模型(LLM)的评估表明,所有三种Agent4RE变体在基于文本的指标上均比领域上下文增强的提示基线平均高出8%。两个增强变体获得最高的LLM作为评判者的评分和人类评分,在四分制上超过两个RE基线约0.8分。这种稳定的性能确立了Agent4RE作为工业环境中实用的端到端RE解决方案的地位。

英文摘要

Existing LLM-based approaches for software Requirements Engineering (RE) typically rely on basic prompting strategies or rudimentary agent collaboration, under-utilizing the full potential of multi-agent systems. Meanwhile, available datasets focus on isolated subtasks, such as requirements extraction, classification, and completeness detection, leaving the absence of an end-to-end RE benchmark that spans from requirements elicitation to generation. We present Agent4RE - a self-refining multi-agent RE system that orchestrates specialized agents and incorporates two iterative improvement loops. To support evaluation, we construct RE-E2E - a real-world dataset built from human-written requirement specifications, enabling end-to-end assessment of RE workflows. Building on this foundation, we further propose two enhanced Agent4RE versions that incorporate either autonomous self-refinement or structured human feedback, and analyze their strengths and limitations across different scenarios. Evaluation on 8 Large Language Models (LLMs) demonstrates that all three Agent4RE variants consistently outperform a domain-context-augmented prompting baseline by average 8% in text-based metrics. The two enhanced variants achieve the highest LLM-as-a-judge and human ratings, surpassing two RE baselines by approximately 0.8 points on a four-point scale. This consistent performance establishes Agent4RE as a practical end-to-end RE solution for industrial environments.

CommentsAccepted to the ASE@POVC track; The E2E requirements engineering benchmark is available https://zenodo.org/records/22815961

DOI:10.1145/3843779.3844634

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑