欺骗增量:基于LLM的智能合约字节码取证对抗性评估
The Deception Delta: Adversarial Evaluation of LLM-Based Smart Contract Bytecode Forensics
- Landeskriminalamt Baden-Württemberg(巴登-符腾堡州刑事调查局)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究对抗性评估22个前沿LLM在智能合约字节码取证中的鲁棒性,发现欺骗可使检测率降低20个百分点,且模型存在合理化现象,当前LLM不能作为独立取证工具。
AI中文摘要:
大语言模型越来越多地被用于区块链取证调查中,以解释未经验证的智能合约字节码。然而,其鲁棒性尚未针对为误导分析而对抗性设计的合约进行系统性测试。我们在13个专门构建的合约(9个欺骗向量,4个对照组)上,跨6种提示策略评估了22个前沿模型,产生了8,528次可分析的非拒绝运行,这些运行针对具有EVM验证真实标签的合约。一个经过校准的LLM-as-judge流水线,辅以两个独立于评判者的指标和50个人类金标准标签,表明相对于功能匹配的对照组,对抗性欺骗使资金流失检测率降低了20.0个百分点(95%置信区间:[17.2, 22.8])。通过多跳调用链、XOR掩码选择器和存储加载的流失参数进行的结构伪装,在几乎所有模型中都抵抗了检测。除了未检测之外,我们还识别出合理化现象:模型正确描述了隐藏的流失机制,但接受了合约的欺骗性框架并将其视为良性,从而产生了积极但错误的关于安全性的证据。简单的防护指令没有提供总体收益,并在两个方向上使个别模型不稳定。结构欺骗在测试的六种提示策略中基本不敏感,这更符合能力限制而非简单的提示问题。只有来自两个提供商的五个模型超过了50%的检测率。在我们的一次性、仅原始字节码协议下,当前的LLM不是可靠的独立取证工具。我们的核心主张不扩展到多轮、工具增强、源码感知或包含反编译器的流水线;源码边界检查作为明确的子集分析报告,而非主要评估的一部分。
英文摘要:
Large language models are increasingly used in blockchain forensic investigations to interpret unverified smart contract bytecode. Their robustness has not been systematically tested against contracts adversarially designed to mislead analysis. We evaluate 22 frontier models on 13 purpose-built contracts (9 deception vectors, 4 controls) across six prompt strategies, yielding 8,528 analyzable non-refusal runs against contracts with documented ground truth. A calibrated LLM-as-judge pipeline, supported by two judge-independent metrics and 50 human gold-standard labels, shows that adversarial deception reduces drain detection by 20.0 percentage points (95% CI: [17.2, 22.8]) relative to functionally matched controls. Structural camouflage via multi-hop call chains, XOR-masked selectors, and storage-loaded drain parameters resists detection across nearly all models. Beyond non-detection, we identify rationalization: models correctly describe the hidden drain mechanism but accept the contract's deceptive framing and dismiss it as benign, yielding positive but incorrect evidence of safety. Simple guard instructions provide no aggregate benefit and destabilize individual models in both directions. Structural deception is largely insensitive across the six tested prompt strategies, more consistent with a capability limitation than with a simple prompting problem. Only five models from two providers exceed 50% detection. Under our single-shot, raw-bytecode-only protocol, current LLMs are not reliable standalone forensic tools. Our central claim does not extend to multi-turn, tool-augmented, source-aware, or decompiler-in-the-loop workflows; a source-code boundary check is reported as an explicit subset analysis rather than as part of the main evaluation.