通过受人工测试启发的工作流程进行多智能体大语言模型协作以生成单元测试
Multi-Agent LLM Collaboration for Unit Test Generation via Human-Testing-Inspired Workflows
浏览论文内容
中文总结 AI 辅助
研究针对现有基于大语言模型的单元测试生成方法的局限,提出TestAgent,通过多智能体协作机制模拟人工测试,设计专门智能体、配备工具API并构建知识图谱,在Java项目实验中取得优异结果,性能优于基线和搜索工具。
中文摘要 AI 辅助
最近,大语言模型(LLMs)的出现促使对自动单元测试生成的研究激增,取得了令人印象深刻的性能并减少了人工工作量。然而,现有的基于LLM的方法仍存在两个主要限制:一是遵循僵化的程序工作流程,未充分利用LLMs的自主推理潜力;二是依赖非针对测试生成定制的基于规则的上下文提取。本文提出TestAgent,一种基于LLM的测试生成方法,通过多智能体协作机制模拟人工测试实践来解决上述限制。它设计了三个专门的智能体,配备工具API并构建测试专用知识图谱。实验结果表明,TestAgent在六个Java项目上取得了高执行率、行覆盖率、分支覆盖率和突变分数,优于基于LLM的基线和基于搜索的工具。
英文摘要
Recently, the emergence of Large Language Models (LLMs) has spurred a surge of research into automated unit test generation, yielding impressive performance and reducing manual effort. However, existing LLM-based approaches still suffer from two major limitations: (1) they follow rigid, procedural workflows that underutilize the autonomous reasoning potential of LLMs, making it difficult to dynamically adapt testing strategies based on real-time feedback; and (2) they rely on rule-based context extraction that is not tailored to test generation, failing to capture fine-grained code dependencies and test-specific knowledge required for deriving test requirements. In this paper, we propose TestAgent, an LLM-based test generation approach that addresses the above limitations by emulating human testing practices via a multi-agent collaboration mechanism. Particularly, TestAgent designs three specialized agents, namely a requirement planner, a test generator, and a test reviewer, to simulate how developers understand, construct, and validate unit tests. To unleash the autonomous capabilities of LLMs, we equip TestAgent with a set of tool APIs that can be invoked dynamically in an on-demand and adaptive manner. To further support repository-level reasoning, TestAgent constructs a test-specialized knowledge graph via static analysis, which captures code entities and their dependencies across the project and persistently stores testing artifacts (e.g., test reports and failure analyses) produced during generation. Experimental results show that TestAgent achieves 97.46% execution rate, 92.34% line coverage, 90.24% branch coverage, and 83.69% mutation score on six Java projects, outperforming LLM-based baselines across all metrics and achieving substantially higher mutation scores than search-based tools.