基于GNN的智能合约漏洞检测中的预模型表示失败
Pre-Model Representation Failures in GNN-Based Smart Contract Vulnerability Detection
浏览论文内容
中文总结 AI 辅助
本文分析了基于GNN的智能合约漏洞检测器的预模型表示层,发现GNNSCVulDetector存在四类表示失败,导致图无法捕获语义并引发误分类,当前准确率未暴露这些问题,其现实普遍性尚待研究。
中文摘要 AI 辅助
本文对基于GNN的智能合约漏洞检测器底层的表示层进行了失败分析。这些系统在任何学习发生之前会将源代码转换为图;如果图未能捕获代码的语义,任何模型改进都无法弥补。我们研究了GNNSCVulDetector并识别出四种失败情况。第一,结构不同的合约会生成逐字节完全相同的图,构成了一种具体的规避攻击。第二,图的构建由一个硬编码的47项变量白名单(包含一个重复项)控制,这限制了提取器可识别的内容。因此,具有不同变量名的相同漏洞会产生不一致的图,图质量会随着命名与白名单的偏离而下降,当没有匹配项时,管道会生成不基于源变量的结构输出。第三,C节点(代表触发重入攻击的外部调用者的图元素)甚至在文献中最典型的易受攻击合约中也不存在。第四,一项受控实验证实这是直接误分类:一个完全可利用的合约被标记为安全,因为从未构建C→W边。所有四种失败情况均通过实验得到验证。文献中当前的准确率是在未暴露这些失败情况的条件下测量的。我们展示了一个由表示层失败直接导致的已确认误分类案例;此类失败在现实世界合约群体中的普遍性仍是一个未解决的实证问题。
英文摘要
This paper is a failure analysis of the representation layer underlying GNN-based smart contract vulnerability detectors. These systems convert source code into graphs before any learning takes place; if the graph fails to capture the code's semantics, no model improvement can compensate. We investigate GNNSCVulDetector and identify four failures. First, structurally different contracts produce byte-for-byte identical graphs, constituting a concrete evasion attack. Second, graph construction is governed by a hardcoded 47-entry variable whitelist (including one duplicate entry), which constrains what the extractor can recognise. As a consequence, identical vulnerabilities with different variable names produce inconsistent graphs, graph quality degrades as naming diverges from the whitelist, and when no entry matches the pipeline produces structural output not grounded in source variables. Third, the C node (the graph element representing the external caller that triggers a reentrancy attack) is absent from even the most canonical vulnerable contract in the literature. Fourth, a controlled experiment confirms this as a direct misclassification: a fully exploitable contract is labelled safe because the C -> W edge is never constructed. All four failures are demonstrated experimentally. Current accuracy figures in the literature are measured under conditions that do not expose these failures. We demonstrate one confirmed case of misclassification caused directly by a representation-layer failure; the prevalence of such failures in real-world contract populations remains an open empirical question.