arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型作为网络-物理系统的反例搜索器

Large Language Models as Falsifiers for Cyber-Physical Systems

Ali ArjomandBigdeli, Jiawei Zhou, Stanley Bak

arXiv 2609.20752首次发表:更新:

发表机构

Stony Brook University(石溪大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文提出LLM-Falsifier,利用大型语言模型最小化STL鲁棒性,通过引入语义信息实现更高效的CPS规范反例搜索,在ARCH-COMP基准上超越现有工具。

AI 中文摘要

反例搜索(Falsification)在网络-物理系统(CPS)中寻找形式化规范的违反示例。当规范以信号时序逻辑(STL)编写时,反例搜索可以被表述为一个鲁棒性优化问题,传统上通过黑盒搜索算法来解决。与此同时,大型语言模型(LLMs)最近在与迭代提示相结合时,展现出令人惊讶的有效优化器能力。在这项工作中,我们连接了这些想法,并引入了LLM-Falsifier,一种基于LLM的方法,通过最小化STL鲁棒性程度来反例搜索规范。除了通用的基于提示的优化之外,我们的关键思想是向LLM暴露对语言模型自然但对标准数值优化器缺失的语义信息,包括自然语言的输入和输出名称、输出轨迹以及最小鲁棒性值的关键时间见证。这些添加使得更智能且更样本高效的鲁棒性搜索成为可能。在ARCH-COMP反例搜索基准测试中,LLM-Falsifier被证明在21个规范中的14个上,以找到反例所需的平均模拟次数衡量,优于基于多种优化范式的现有反例搜索工具,从基于代理和贝叶斯优化到基于搜索的测试。

英文摘要

Falsification searches for counterexamples to formal specifications in cyber-physical systems (CPS). With specifications written in Signal Temporal Logic (STL), falsification can be formulated as a robustness optimization problem, traditionally tackled with black-box search algorithms. In parallel, large language models (LLMs) have recently emerged as surprisingly effective optimizers when coupled with iterative prompting. In this work, we connect these ideas and introduce LLM-Falsifier, an LLM-based approach that falsifies specifications by minimizing the STL robustness degree. Beyond generic prompt-based optimization, our key idea is to expose the LLM to semantic information that is natural for language models but absent from standard numerical optimizers, including natural-language input and output names, output trajectories, and critical-time witnesses for the minimum robustness value. These additions enable smarter and more sample-efficient robustness search. On the ARCH-COMP falsification benchmarks, LLM-Falsifier is shown to outperform existing falsification tools based on a range of optimization paradigms, from surrogate-based and Bayesian optimization to search-based testing, on 14 of 21 specifications when measured by the average number of simulations required to find a counterexample.

Comments22 pages, 5 figures, 3 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑