arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Tangent:基于大语言模型的智能体应用测试实践的实证研究

Tangent: An Empirical Study of Testing Practices for LLM-Based Agent Applications

Rangeet Pan, Tyler Stennett, Divya Sankar, Bridget McGinn, Alessandro Orso, Raju Pavuluri, Saurabh Sinha, Maja Vukovic

arXiv 2608.08413首次发表:更新:

AI 中文总结

Tangent通过对开源项目和行业从业者的研究,分析了LLM智能体应用的测试现状,发现其以单元测试为主、存在覆盖不足等问题,并提出了相关研究方向。

AI 中文摘要

基于大语言模型(LLM)构建的智能体越来越多地被用于开发执行复杂多步骤任务的应用,这些任务涉及推理、工具使用以及与外部环境的交互。尽管LLM智能体的基准测试取得了快速进展,但很少有研究试图了解此类系统在实践中是如何测试的。特别是,智能体应用的测试级别、目标、数据模式、测试复杂度和验证策略仍未得到充分探索。在本文中,我们使用从开源项目中挖掘的大型语料库,对LLM智能体应用的测试实践进行了实证研究。我们构建了一个包含智能体应用、工具和测试的大规模数据集,并对来自240个模块的2572种测试方法进行了手动标注。通过分析,我们得出了涵盖测试夹具、数据、目标和断言的23种测试模式的分类法,并按级别(单元、模块、集成)对测试进行了表征。我们还补充了对10名构建智能体系统的行业资深从业者的结构化访谈。我们的结果显示,LLM智能体应用的测试以范围狭窄的单元测试为主,对复杂交互、现实场景和非功能需求的覆盖有限。测试经常依赖简单的输入、大量的模拟和浅层验证,且与智能体相关的测试表现出较低的结构复杂性。虽然行业实践比开源项目更强调非功能测试,但两者都存在共同的差距,包括缺乏正式的测试基础、测试目标不明确以及生成高质量测试数据的挑战。基于这些发现,我们概述了迈向更系统、更严格的智能体应用测试的研究方向,包括智能体可测试性的基础、形式化的测试目标以及基于故障的测试技术。

英文摘要

Agents built on large language models (LLMs) are increasingly used to build applications that perform complex, multi-step tasks involving reasoning, tool use, and interaction with external environments. Despite rapid progress in benchmarking LLM-based agents, very few studies have attempted to understand how such systems are tested in practice. In particular, testing levels, objectives, data patterns, test complexity, and validation strategies for agent applications remain underexplored. In this paper, we present an empirical study of testing practices in LLM-based agent applications using a large corpus of mined open-source projects. We construct a large-scale dataset of agent applications, tools, and tests, and manually label 2,572 test methods from 240 modules. From this analysis, we derive a taxonomy of 23 testing patterns across test fixtures, data, objectives, and assertions, and characterize tests by level (unit, module, integration). We complement this with structured interviews of 10 senior industry practitioners building agentic systems. Our results show that testing of LLM-based agent applications is dominated by narrowly scoped unit tests, with limited coverage of complex interactions, realistic scenarios, and non-functional requirements. Tests frequently rely on simplistic inputs, heavy mocking, and shallow validation, and agent-related tests exhibit low structural complexity. While industry practice places greater emphasis on non-functional testing than open-source projects, both reveal common gaps, including the lack of formal testing foundations, unclear test objectives, and challenges in generating high-quality test data. Based on these findings, we outline research directions toward more systematic and rigorous testing of agent applications, including foundations for agent testability, formalized test objectives, and fault-based testing techniques.

CommentsAccepted at ASE'26

DOI:10.1145/3832783.3837414

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑