AI 中文总结
研究针对工业智能体基准测试场景有限问题,扩展AssetOpsBench并引入ScenarioGeneratorAgent管道,通过多种优化方式提高可扩展性,实现智能电网变压器场景生成的高效优化,有效扩展工业智能体基准测试且保证场景质量。
AI 中文摘要
工业智能体基准测试需要整合遥测、故障模式、维护记录和领域标准的现实评估场景。现有基准测试依赖人工编写场景且覆盖资产类别有限。我们扩展了AssetOpsBench,加入智能电网变压器资产类别和四种基于IEC的诊断工具。还引入ScenarioGeneratorAgent合成工业智能体场景生成管道,通过多种方式优化以提高可扩展性。在智能电网变压器场景生成中,优化使50个场景的端到端运行时间缩短8倍且保持质量,结果表明基于标准的合成场景生成可有效扩展工业智能体基准测试且不牺牲场景质量。
英文摘要
Industrial agent benchmarks require realistic evaluation scenarios that integrate telemetry, failure modes, maintenance records, and domain standards. However, existing benchmarks such as AssetOpsBench rely on manually authored scenarios and cover a limited set of asset classes. We extend AssetOpsBench with a Smart Grid Transformer asset class and four IEC-grounded diagnostic tools for health-index prediction, dissolved-gas analysis, winding-temperature assessment, and load-profile assessment. We further introduce ScenarioGeneratorAgent, a pipeline for synthetic industrial-agent scenario generation. The pipeline constructs evidence-grounded asset profiles, allocates coverage-aware scenario budgets across operational domains, and generates candidates through a hybrid validation-and-repair loop that enforces schema validity, tool reachability, physical plausibility, standards alignment, and deduplication. To improve scalability, we apply two-level caching, parallel focus-group generation, thread-pool offloading, batched LLM calls, and early rejection filtering. On Smart Grid Transformer scenario generation, these optimizations reduce end-to-end runtime by $8\times$ for 50 scenarios while preserving quality, achieving a composite quality score of $74.2 \pm 1.9$ compared with $73.8 \pm 3.0$ for the unoptimized baseline. These results show that standards-grounded synthetic scenario generation can efficiently expand industrial-agent benchmarks without sacrificing scenario quality.
Comments19 pages, 3 appendices