arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.17546cs.SE

基于经验证的LLM推断依赖关系与响应驱动优化的REST API测试

REST API Testing with Verified LLM-Inferred Dependencies and Response-Driven Refinement

Tu Nguyen, Thanh Nguyen, Huy Nguyen, Viet Nguyen, Tien N. Nguyen, Vu Nguyen

AI总结:

本文提出APIPilot框架,通过执行验证LLM推断的REST API依赖关系,结合响应驱动优化,在16个真实API上实现高覆盖率与成功率,优于现有基线。

AI中文摘要:

测试RESTful API需要生成满足操作、参数及运行时创建资源之间依赖关系的API调用序列。近期基于LLM的方法可从OpenAPI规范中推断此类依赖关系并生成测试序列,但通常将LLM推断的关系视为正确而不进行基于执行的验证,这会引入虚假依赖、遗漏可行操作链并生成不可行的测试。本文提出APIPilot,一种用于REST API测试的经执行验证的框架。APIPilot首先使用结构启发式方法和基于LLM的语义推理从OpenAPI规范中推导候选生产者-消费者依赖关系,随后将这些依赖关系视为假设并通过具体API执行对其进行验证,再将其用于测试生成。经验证的依赖关系被组织成依赖图,APIPilot通过有界top-k图遍历从该图构建感知覆盖率的工作流,将语义依赖推断与序列构建分离。为改进后续测试,APIPilot还执行响应驱动优化:分析运行时响应以更新资源池、调整输入生成约束并修剪或修正无效依赖映射。对16个真实REST API服务的实证评估显示,APIPilot实现了92.3%的操作覆盖率、最高58.6%的代码覆盖率和88.1%的工作流执行成功率,优于基于LLM的和传统的REST API测试基线。APIPilot还检测到197个独特的5xx错误及规范-执行不匹配,证明将依赖推断建立在执行反馈基础上的益处。

英文摘要:

Testing RESTful APIs requires generating sequences of API calls that satisfy dependencies among operations, parameters, and runtime-created resources. Recent LLM-based approaches infer such dependencies and generate test sequences from OpenAPI specifications, but they often treat LLM-inferred relationships as correct without execution-based validation. This can introduce spurious dependencies, miss feasible operation chains, and produce infeasible tests. In this paper, we propose APIPilot}, an execution-validated framework for REST API testing. APIPilot first derives candidate producer-consumer dependencies from OpenAPI specifications using structural heuristics and LLM-based semantic reasoning. It then treats these dependencies as hypotheses and validates them through concrete API executions before using them for test generation. The validated dependencies are organized into a dependency graph from which APIPilot constructs coverage-aware workflows via bounded top-k graph traversal, separating semantic dependency inference from sequence construction. To improve subsequent tests, APIPilot further performs response-driven refinement: runtime responses are analyzed to update resource pools, adjust input-generation constraints, and prune or revise invalid dependency mappings. Empirical evaluation on 16 real-world REST API services shows that APIPilot achieves 92.3% operation coverage, up to 58.6% code coverage, and an 88.1% workflow execution success rate, outperforming both LLM-based and traditional REST API testing baselines. APIPilot also detects 197 unique 5xx failures and specification-execution mismatches, demonstrating the benefit of grounding dependency inference in execution feedback.

补充信息

↑