arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

Mind the Gap: 检测 DAO 治理中的描述-执行不匹配攻击

Mind the Gap: Detecting Description-Execution Mismatch Attacks in DAO Governance

Bowen Cai, Nanzi Yang, Weiheng Bai, Youshui Lu, Yajin Zhou, Kangjie Lu

arXiv 2609.13601首次发表:更新:

发表机构

University of Minnesota; Old Dominion University; Xi’an Jiaotong University; The Chinese University of Hong Kong(明尼苏达大学; 老道明大学; 西安交通大学; 香港中文大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对 DAO 治理中描述与代码执行不匹配的欺骗性提案攻击,提出基于 DAO 无关模拟和 LLM 证据映射的检测框架,在真实以太坊数据上实现高精确率与召回率,并具备跨模型鲁棒性。

AI 中文摘要

去中心化自治组织(DAO)通过基于提案的流程来更改协议:发起者提交提案,成员根据其自然语言描述进行投票,若提案通过,项目则执行其背后的代码。这一过程本质上容易受到欺骗性提案的攻击,即描述意图与实际代码执行不匹配。恶意提案者可以提交看似良性的描述以通过投票,而实际执行的代码却转移资金或夺取协议控制权,我们称之为描述-执行不匹配(DEMI)攻击。我们提出了首个在真实 DAO 治理中检测 DEMI 的系统性框架。首先,DAO 部署的多样性使得统一、可扩展的分析变得困难;我们通过一个与 DAO 无关的模拟框架解决了这一问题,该框架从历史链上交易中构建每个 DAO 的治理画像,然后驱动每个新提案经历完整的治理生命周期以获得其执行行为。其次,自由形式的描述和结构化的执行轨迹难以比较;我们通过一种证据映射范式解决了这一问题,该范式要求大语言模型(LLM)定位每个操作的明确文本依据,而非给出整体判断,从而显著提高了相对于直接查询的精确率和召回率。在真实世界以太坊 DAO 治理的大规模数据集上,我们的模拟为 92.7% 的活跃 DAO 和 89.3% 的已执行提案推导出了执行结果,远超现有平台。在分层交叉验证下,我们的检测器达到了 81.7% 的平均精确率和 98.3% 的平均召回率,且证据映射可跨 LLM 供应商泛化,而非依赖于单一模型。在系统性的红队/蓝队评估中,即使面对知晓检测器的对手,鲁棒性守卫也能防御大多数自适应攻击,且这种鲁棒性可泛化到留出的提案上。

英文摘要

Decentralized autonomous organizations (DAOs) change protocols through a proposal-based process: initiators submit a proposal, members vote on it based on its natural-language description, and if it passes, the project executes the code behind it. This process is inherently vulnerable to deceptive proposals, where the described intent and the actual code execution mismatch. A malicious proposer can submit a benign-looking description to pass voting while the executed code transfers funds or seizes control of the protocol, which we call a Description-Execution Mismatch (DEMI) attack. We present the first systematic framework for DEMI detection in real DAO governance. First, the diversity of DAO deployments makes a unified, scalable analysis difficult; we address this with a DAO-agnostic simulation framework that builds a per-DAO governance profile from historical on-chain transactions and then drives each new proposal through the full governance lifecycle to obtain its execution behavior. Second, free-form descriptions and structured execution traces are hard to compare; we address this with an evidence-mapping paradigm that requires an LLM to locate explicit per-action textual justifications rather than issue a holistic judgment, substantially improving precision and recall over direct querying. On a large-scale dataset of real-world Ethereum DAO governance, our simulation derives execution results for 92.7% of active DAOs and 89.3% of executed proposals, far exceeding existing platforms. Our detector reaches 81.7% mean precision and 98.3% mean recall under stratified cross-validation, and evidence mapping generalizes across LLM vendors rather than depending on one model. Under a systematic red-team/blue-team evaluation, the Robustness Guard defends most adaptive attacks even against an adversary that knows the detector, and this robustness generalizes to held-out proposals.

CommentsExtended version of the paper accepted at ACM CCS 2026. 18 pages, 11 figures, 9 tables; includes the full technical appendix

DOI:10.1145/3830454.3846550

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑