RFCLLM:评估大语言模型对网络协议状态机推理能力
RFCLLM: Evaluating LLMs' Reasoning Ability of Network Protocol State Machines
- Northeastern University(东北大学)
- Purdue University(普渡大学)
机构由 AI 辅助整理,请以论文原文为准。
中文总结 AI 辅助
本文通过4项任务和1482个查询评估大语言模型对协议状态机的推理能力,发现其隐式表示与地面真值模型存在偏差,为验证其在FSM推理中的可信度提供依据。
中文摘要 AI 辅助
将文本规范映射为形式化表示对于确保协议设计和实现的正确性至关重要。由大语言模型生成的映射,用于网络安全或测试时,被假定为对规范具有完美的理解,但这在实践中可能并不成立。本文的目标是评估大语言模型能在多大程度上正确解释规范。我们考察了大语言模型对通过自然语言描述的有限状态转移系统的隐式表示,与人工生成的地面真值模型之间的对齐程度。我们为16个协议设计了4项任务和1482个任务查询。我们评估了不同的评判者偏差,观察了任务之间固有的难度差距,研究了4种上下文类型的影响,以及协议特征的影响。我们的工作朝着验证大语言模型在协议规范的有限状态机(FSM)推理中是否真正可信迈出了一步。
英文摘要
Mapping textual specifications into formal representations is essential for ensuring the correctness of protocol designs and implementations. LLM-generated mappings, used for networking security or testing, are assumed to capture a perfect understanding of the specification, which may not hold in practice. The goal of this paper is to assess the extent to which LLMs can interpret the specification correctly. We examine the degree to which an LLM's implicit representation of a finite-state transition system-defined via natural language descriptions-aligns with a manually generated ground-truth model. We designed 4 tasks and 1482 task queries for 16 protocols. We evaluated different judge biases, observed the inherent difficulty gaps between tasks, looked into the effect of 4 context types, and the influence of protocol characteristics. Our work contributes to a step toward verifying whether LLMs can really be trusted in FSM (Finite State Machine) reasoning of protocol specifications.