发表机构
School of Systems Science and Industrial Engineering, Binghamton University, State University of New York; Meta Platforms, Inc.(系统科学与工业工程学院,宾夕法尼亚州立大学布林茅尔分校; Meta平台公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对大规模人工智能数据中心平台硬件验证测试计划手动编写的问题,提出基于生成式人工智能多智能体架构,依据自愈验证文档和物料清单自动生成测试计划,提高了覆盖范围、编写效率和可追溯性。
AI 中文摘要
大规模人工智能数据中心平台包含数千个异构硬件组件,其验证需要全面的故障注入测试计划。目前这些计划是手动编写的,过程繁琐、易出错且依赖机构知识。本文提出一种生成式人工智能多智能体架构,能根据自愈验证文档和组件物料清单这两个规范输入自动生成结构化硬件验证测试计划。通过摄取代理、分类代理和生成代理的协同工作,输出符合标准化模式,可直接导入内部验证软件。在两个生产平台上与手动基线相比,该框架实现了覆盖范围的大幅扩展,将编写时间从数天缩短至数小时,具有高度的可追溯性和跨平台的可移植性。自动化和专家评估证实了其高提取保真度和对新场景的高接受度。
英文摘要
Large-scale AI datacenter platforms comprise thousands of heterogeneous hardware components whose validation requires comprehensive fault injection test plans. Today these plans are authored manually: engineers review hardware self-healing validation documents and bills of materials, enumerate failure modes per field-replaceable unit, and produce flat lists of single-layer test cases. This process is labor-intensive, error-prone, and dependent on institutional knowledge; coverage gaps surface late, traceability to source specifications is implicit, and the effort is largely repeated per platform. This paper presents a generative AI multi-agent architecture that automates the generation of structured hardware validation test plans from two canonical inputs: self-healing validation documents, which enumerate known failure modes and their detection and remediation behaviors per field-replaceable unit, and component Bills of Material. An ingestion agent normalizes heterogeneous inputs into a canonical representation; a classification agent maps components to functional domains via contextual reasoning over part descriptions and sub-category hierarchies; and a generation agent synthesizes test cases by combining normalized failure modes with domain-classified data, filling gaps and producing edge cases. The output conforms to a standardized schema for direct import into internal validation software. Evaluated on two production platforms against manual baselines, the framework achieves coverage expansions of 74.2% and 51.4%, cutting authoring from days to hours. It yields fully traceable mappings from each test case to its source specification, and its multi-agent decomposition is portable across platform generations. Automated and expert evaluations confirm 100% extraction fidelity and high acceptance of new scenarios, validating the framework as a robust human-in-the-loop force multiplier.