复杂智能体,浅显测试:揭示并增强野外智能体框架的测试充分性
Complex Agents, Shallow Tests: Demystifying and Enhancing Test Adequacy of Agent Harness in the Wild
- University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
- University of California, Berkeley(加州大学伯克利分校)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
针对智能体框架测试不足的问题,提出首个面向框架的测试生成技术HarnessTester,通过契约支持提升LDH代码覆盖率,显著优于现有方法,并发现大量真实缺陷。
AI中文摘要:
基于大语言模型的智能体系统正成为一种新的软件范式。现代智能体通常由骨干大语言模型和环绕的框架(harness)组成,该框架作为智能体执行的操作软件基础设施。随着智能体框架日益复杂,智能体遭受多种框架实现缺陷,引发了显著的可靠性担忧。在本工作中,我们进行了首个实证研究,系统性地调查了真实世界智能体系统中框架的测试充分性。我们的分析揭示,智能体框架仍存在大量测试不足。特别是,依赖大语言模型的框架(LDH)代码,尽管在处理大语言模型输出和控制智能体行为方面起关键作用,却受到有限的测试关注,其行和分支的现有测试覆盖率不足一半。受这些发现的启发,我们进一步提出了HarnessTester,这是首个面向框架的测试生成技术,它引入显式的智能体-框架契约支持,以构建符合契约的测试设置,并广泛执行LDH代码。我们的评估显示,HarnessTester在行/分支覆盖率提升上分别比最先进的通用测试生成技术高出75.95%/84.76%,在变异分数提升上高出69.89%。此外,HarnessTester在广泛使用的智能体系统(如OpenClaw)中检测到122个真实世界框架缺陷,其中88个是先前未知的缺陷,69个已得到智能体开发者的确认。这些结果凸显了HarnessTester在提高测试充分性和保障真实世界智能体系统可靠性方面的实际有效性。
英文摘要:
LLM-based agentic systems are emerging as a new software paradigm. Modern agents are typically composed of backbone LLMs and a surrounding harness that serves as the operational software infrastructure for agent execution. As agent harnesses grow increasingly complex, agents suffer from diverse harness implementation bugs, raising substantial reliability concerns. In this work, we conduct the first empirical study to systematically investigate the test adequacy of harness in real-world agentic systems. Our analysis reveals that agent harness remains substantially undertested. In particular, LLM-dependent harness (LDH) code, despite its critical role in processing LLM outputs and governing agent behavior, receives limited testing attention, with less than half of its lines and branches covered by existing tests. Motivated by these findings, we further propose HarnessTester, the first harness-oriented test generation technique that incorporates explicit agent-harness contract support to construct contract-faithful test setups and extensively exercise LDH code. Our evaluation shows that HarnessTester substantially outperforms state-of-the-art general-purpose test generation techniques in achieving 75.95%/84.76% larger line/branch coverage gains and 69.89% larger mutation-score gains. Furthermore, HarnessTester detects 122 real-world harness bugs in widely-used agentic systems (e.g., OpenClaw), among which, 88 bugs are previously-unknown bugs and 69 bugs have been confirmed by agent developers. These results highlight the practical effectiveness of HarnessTester in improving test adequacy and assuring the reliability of real-world agentic systems.