发表机构
Systems Engineering, King Fahd University of Petroleum; Faculty of Computer; Information Systems Islamic University of Madinah Madinah, Saudi Arabia; Division of Information Science, Nara Institute of Science(; ; ; )
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对人工智能和网络物理系统中高级需求与低级测试脱节问题,提出VNVSpec开源框架,使V&V规范机器可读可执行,能检查需求质量、链接需求与测试结果等,通过自我应用评估并可扩展用于相关测试。
AI 中文摘要
现代软件团队有用于低级测试的成熟工具,如pytest、JUnit和Jest,编写单元测试并在每次提交时运行成本低。系统工程也为设计验证与确认(V&V)制定了严格原则。但实际中二者很少连接。对于人工智能和网络物理系统,这种差距成本越来越高。我们引入VNVSpec框架,使V&V规范机器可读且可执行。用户可直接陈述或从标准目录导入高级需求,框架能检查需求质量、支持分解、链接需求与测试结果并生成报告。通过自我应用评估框架,还讨论了其对黑盒人工智能模型和人工智能编码代理测试的扩展。框架及其测试套件、目录和基准脚本可在指定网址获取。
英文摘要
Modern software teams have mature tools for low-level testing, such as pytest, JUnit, and Jest, which make it inexpensive to write unit tests and run them on every commit. Systems engineering, in parallel, has developed rigorous principles for design verification and validation (V&V), which has worked very well across engineering discipline to align user expecations and requirements with developers' deliverables. In practice, however, the two rarely connect, and the link between users' high-level requirements and the low-level tests that machines actually run is maintained by hand, if at all. This gap is increasingly costly for AI-enabled and cyber-physical systems, for which regulators now ask for traceable evidence that high-level requirements are met, while raw test results provide little of the structure such evidence requires. We introduce VNVSpec, an open-source framework that makes V&V specifications machine-readable and executable. With this framework, users state high-level requirements directly or import them from catalogs derived from published standards. Then, the framework checks requirement quality, supports decomposition into module-level requirements with explicit metrics and acceptance criteria, links these requirements to test results through a traceability graph, and compiles the collected evidence into verdicts and audit-ready reports. We evaluate the framework by self-application, in which it is continuously assessed in CI against its own specification of 36 requirements verified by 449 tests, completed within limited time which scales linearly and thus can handle up to 10,000 requirements. We also discuss how the framework extends to testing black-box AI models and AI coding agents. The framework, its full test suite, the catalogs, and the benchmark scripts are available at https://github.com/ai-vnv/vnvspec.