发表机构
International Business Machines (IBM); Global Atlantic Financial; Docusign; Salesforce Inc; The Home Depot(国际商业机器公司(IBM); 全球大西洋金融公司; DocuSign公司; Salesforce公司; 家得宝公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对现代软件质量保证需求,提出AINTMA多智能体AI系统。通过六个专门智能体及安全通信框架协调,结合生成式智能等技术。经实验评估,该系统在测试优先级、周期时间、缺陷逃逸率等方面表现出色,推进了云环境下软件质量管理。
AI 中文摘要
现代软件质量保证需要能在分布式云环境中进行自适应决策的智能自主系统。本文提出了AINTMA(智能测试管理架构),这是一个多智能体AI系统,将传统测试管理转变为自主质量智能生态系统。AINTMA部署了六个专门的AI智能体,通过安全的多智能体通信框架在云原生微服务基础设施上进行协调。生成式质量智能智能体使用大语言模型生成自然语言质量叙述、缺陷风险摘要和数据增强测试建议。RL优先级智能体将测试选择建模为马尔可夫决策过程,从大规模历史测试执行数据中学习上下文策略。通过具有OAuth2/JWT认证、加密智能体间消息传递和多租户隔离的零信任API网关来加强安全云通信。对12个异构软件项目进行18个月的评估表明,测试优先级准确率达88.4%,测试周期时间减少43%,缺陷逃逸率从8.3%降至2.1%,9个月回收期的投资回报率为340%。智能体架构可扩展到50000多个测试用例,响应时间低于400毫秒,生成式智能模块的开发者有用性评级为4.3/5.0。AINTMA表明,结合自主多智能体协调、生成式智能和安全智能连接的智能AI可以从根本上推进云规模企业环境中的软件质量管理。
英文摘要
Modern software quality assurance demands intelligent, autonomous systems capable of adaptive decision-making across distributed cloud environments. This paper presents AINTMA (Agentic Intelligent Test Management Architecture), a multi-agent agentic AI system that transforms traditional test management into an autonomous quality intelligence ecosystem. AINTMA deploys six specialized AI agents (Test Discovery, Risk Assessment, Reinforcement Learning Prioritization, Execution Orchestration, Generative Quality Intelligence, and Cloud Security Monitor) coordinated through a secure multi-agent communication framework over a cloud-native microservices infrastructure. The Generative Quality Intelligence agent employs large language models to produce plain language quality narratives, defect risk summaries, and data-augmented test recommendations. The RL Prioritization agent models test selection as a Markov Decision Process, learning contextual policies from large-scale historical test execution data (47 features, rolling 36-month window). Secure cloud communication is enforced through a zero-trust API gateway with OAuth2/JWT authentication, encrypted inter-agent messaging, and multi-tenant isolation. Evaluation across 12 heterogeneous software projects over 18 months demonstrates: 88.4% test prioritization accuracy (APFD, vs. 51.2% random, 82.1% best commercial baseline); 43% test cycle time reduction; defect escape rate reduced from 8.3% to 2.1%; 340% ROI at 9-month payback. The agentic architecture scales to 50,000+ test cases with sub-400ms response time, and the generative intelligence module achieves 4.3/5.0 developer usefulness rating. AINTMA demonstrates that agentic AI, combining autonomous multi-agent coordination, generative intelligence and secure smart connectivity, can fundamentally advance software quality management in cloud-scale enterprise environments.
Comments11 pages, 2 figures, 4 tables, Submitted to AICCONS (AIP Conference Proceedings format)