发表机构
Nankai University; Zhongguancun Laboratory(南开大学; 中关村实验室)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究首次对比LLM在模拟与开发模式下的安全表现,发现模拟模式可筛选但不可替代开发测试,且模型能背诵安全要求时仍需书面记录。
AI 中文摘要
大型语言模型(LLM)既被用于模拟网络协议实现(如蜜罐所做的那样),也被用于编写这些实现。先前的工作分别评估了这两种用途,一种是从模型的回答或对话中评估,另一种是从其生成的代码中评估。我们观察到,知识探针(询问模型实现需要哪些安全检查)和对话可以认可生成程序所缺乏的安全检查,但没有研究将这两者与同一模型编写的代码进行比较。为了填补这一空白,我们首次在模拟模式(S模式,模型在对话中扮演实现)和开发模式(D模式,模型编写程序且程序受到攻击)之间进行了此类比较,并以知识探针作为基线。我们设计并实现了一个测试平台,用一个攻击套件和一个预言机来评判所有三种模式,并在四个协议实现上评估了15个LLM。我们发现,在完整的安全规范下,中位模型在重组、HTTP和防火墙实现上的攻击成功率在两种模式下均至多为4%,但在DNS上,S模式为13%,D模式为15%。在D模式下,向规范中添加缺失的DNS要求可将攻击成功率从100%降至34%,且仅改变该检查;而删除一条明确规则则会削弱程序在其他检查上的表现。事先询问模型并不能预测程序具有哪些检查,因为14个编写DNS解析器的模型中有13个提到了该检查,但所有42个程序都缺乏它。S模式是一个有用的初步过滤器,能标记出82%的真实DNS漏洞,并让96%的安全案例不被标记。它在两个方向上都会出错:过度报告D模式未实现的安全检查,以及少报D模式已实现的检查。S模式可以筛选但不能替代D模式测试,并且即使模型能够背诵安全要求,也应将其书面记录下来。
英文摘要
LLMs are used both to simulate network protocol implementations, as honeypots do, and to write them. Prior work evaluates the two uses separately, from a model's answers or conversations in one case and from its generated code in the other. We observe that a knowledge probe (asking the model which security checks an implementation needs) and a conversation can credit security checks that the generated program lacks, but no study has compared them with the code the same model writes. To fill this gap, we present the first such comparison between simulation mode (S-mode), where the model plays the implementation in a conversation, and development mode (D-mode), where the model writes the program and the program is attacked, with a knowledge probe as a baseline. We design and implement a harness that judges all three with one attack suite and one oracle, and evaluate 15 LLMs on four protocol implementations. We find that with the full security specification, the median model's attack success is at most 4\% on reassembly, HTTP, and firewall implementations in both modes, but 13\% in S-mode and 15\% in D-mode on DNS. In D-mode, adding the missing DNS requirement to the specification cuts that attack from 100\% to 34\% and changes only that check, whereas deleting a stated rule weakens the program on other checks too. Asking the model beforehand does not predict which checks a program has, since 13 of the 14 models that write a DNS resolver name the check, yet all 42 programs lack it. S-mode is a useful first filter, flagging 82\% of the real DNS vulnerabilities and leaving 96\% of the safe cases unflagged. It errs in both directions, over-reporting security checks that D-mode does not implement and under-reporting checks that D-mode does. S-mode can screen but not replace D-mode testing, and security requirements should be written down even when a model can recite them.
Comments21 pages, 6 figures