评估用于符号安全协议分析的大语言模型
Evaluating Large Language Models for Symbolic Security Protocol Analysis
浏览论文内容
中文总结 AI 辅助
研究评估大语言模型对符号安全协议的分析能力,通过在130个协议上测试GPT和DeepSeek,发现聊天与推理模型各有优劣,在认证目标上表现差,保密性较好,结果不稳定,整体LLMs无法与形式化验证相比,仅可作预筛选。
中文摘要 AI 辅助
安全协议验证依赖于如ProVerif和OFMC等形式化工具。本研究评估大语言模型(LLMs)是否能进行可比分析。在130个混淆的AnB/AnBx协议(涵盖388个安全目标)上对GPT和DeepSeek进行三轮聊天和推理模式测试,并与ProVerif和OFMC对比评分。聊天模型在精度低于31%时召回率达69%至81%。推理模型则相反,GPT精度达66.5%,DeepSeek达45.4%,但检测到的攻击略超一半。所有模型在认证目标上表现最差,保密性是例外,推理模式下F1高达95.7%。结果在各轮中不稳定,自我报告的信心普遍高但与正确性无显著关联。在该基准上LLMs无法与形式化验证匹配,至多可作为预筛选过滤器。
英文摘要
Security protocols verification relies on formal tools such as ProVerif and OFMC. This study evaluates whether large language models (LLMs) can perform comparable analysis. We test GPT and DeepSeek in chat and reasoning modes over three runs on 130 obfuscated AnB/AnBx protocols covering 388 security goals, scored against ProVerif and OFMC. Each provider uses a single model in both modes, switching reasoning on and off, so both contrasts isolate reasoning itself. Chat models achieve 72.7% recall at 27.3% precision for GPT and 69.3% recall at 27.2% precision for DeepSeek. Reasoning models reverse this trade-off, reaching 66.5% precision and 54.5% recall for GPT and 45.4% precision and 57.3% recall for DeepSeek. Enabling reasoning lifts precision from 27.3% to 64.8% for GPT and from 27.2% to 44.4% for DeepSeek on the consolidated verdict. The goal set is imbalanced, with 89 vulnerable goals against 299 secure ones; a trivial always-secure predictor scores 77.1% accuracy, which only GPT reasoning exceeds. All models perform worst on authentication goals: reasoning models detect well under half of injective and non-injective agreement attacks, whereas chat models over-flag them at low precision. Confidentiality is the exception, with F1 up to 95.7% in reasoning mode. Verdicts are unstable across runs: identical on 89.7% of goals for GPT reasoning, 74.0% for DeepSeek reasoning, 70.1% for GPT chat, and 61.6% for DeepSeek chat. Self-reported confidence is uniformly high yet shows no meaningful correlation with correctness. All results rest on a single zero-shot prompt and two model providers, which limits generalisability. On this benchmark, LLMs do not match formal verification, but may serve, at best, as pre-screening filters.
发表机构
- Teesside University(蒂赛德大学)
机构由 AI 辅助整理,请以论文原文为准。