AI 中文总结
针对RESTful API测试中定义可靠测试预言机的挑战,Restor框架利用强化学习,通过数据增强管道微调大语言模型生成测试断言,在工业数据集和生产工作流程中表现出色,提升测试用例采用率并减少人工QA工作。
AI 中文摘要
现代REST API测试在定义可靠的测试预言机方面面临关键挑战,尤其是在敏捷工业环境中,正式规范常缺失或过时,新部署端点无历史执行日志。本文提出Restor框架,能在黑盒设置下从单个观察到的请求 - 响应对生成可执行测试断言。它利用新颖数据增强管道,通过组相对策略优化微调轻量级大语言模型。在工业数据集上评估,Restor显著优于基线和通用模型,在字节跳动生产CI/CD工作流程中部署,提高了自动生成测试用例的采用率,减少了人工质量保证工作。
英文摘要
Modern REST API testing faces a critical challenge in defining reliable test oracles, particularly in agile industrial environments where formal specifications (e.g., OpenAPI) are frequently missing or outdated, and historical execution logs are unavailable for newly deployed endpoints. In this paper, we present Restor (Reinforcement Enhanced Single-Traffic Oracle generator for REST APIs), a framework that generates executable test assertions from a single observed request-response pair in a black-box setting. Unlike existing approaches that rely on rule-based templates or massive training logs, Restor utilizes a novel data augmentation pipeline to fine-tune a lightweight Large Language Model (LLM) via Group Relative Policy Optimization (GRPO). This training process enables the model to internalize testing "common sense" by optimizing a reward function that jointly encourages: (i) the selection of stable, semantically meaningful fields for validation and the avoidance of dynamic noise (e.g., timestamps or trace IDs); (ii) the generation of robust assertions that withstand logic variations. We evaluate Restor on an industrial dataset comprising over 2,300 API traces across 246 real-world services. Comprehensive experiments demonstrate that Restor significantly outperforms prompt-engineered baselines and generalist models, achieving a superior $F_1$ score of 85.42% in key field identification and increasing the proportion of semantically accurate assertions. Furthermore, deployment in a production CI/CD workflow at ByteDance confirms its practical value: the system raised the adoption rate of automatically generated test cases from 74.1% to over 96%, substantially reducing manual Quality Assurance (QA) effort while ensuring high execution stability.
CommentsAccepted to ISSTA 2026. 22 pages, 8 figures, 2 tables