arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

揭秘智能合约审计中的智能体技能:设计、有效性与行为影响

Demystifying Agent Skills for Smart Contract Auditing: Design, Effectiveness, Behavioral Impact

Cuifeng Gao, Juantao Zhong, Jiachi Chen, Shuai Wang, Daoyuan Wu

arXiv 2609.29454首次发表:更新:

发表机构

Lingnan University; The Hong Kong Polytechnic University; State Key Laboratory of Blockchain and Data Security, Zhejiang University; Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security; The Hong Kong University of Science and Technology(岭南大学; 香港理工大学; 浙江大学区块链与数据安全全国重点实验室; 杭州高新技术开发区(滨江)区块链与数据安全研究院; 香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究系统收集并评估83个智能合约审计技能,揭示其设计异构、有效性取决于模型而非框架,并发现技能触发是瓶颈,Codex/GPT-5.5提升检测得分22.8%。

AI 中文摘要

LLM智能体,尤其是Claude Code和OpenAI Codex,正逐渐成为超越单纯编码智能体的多功能工具。这些智能体可以通过技能(可复用的工件,封装领域知识、工作流程和工具使用指令)得到增强。然而,迄今为止,关于这些技能如何设计,以及它们在实践中如何影响智能体的有效性和行为,我们知之甚少。在本文中,我们在智能合约安全审计这一智能体已展现出巨大潜力的领域,研究了这些问题。我们系统地收集了83个来自实际场景的智能合约审计技能,并在EVMBench上,跨七种智能体-模型配置对其进行了评估。我们的研究考察了三个维度:(i)审计技能的设计特征,包括其结构、知识表示、工作流程和工具依赖;(ii)它们在改进漏洞检测方面的有效性;以及(iii)它们对智能体执行轨迹的影响。我们发现,审计技能大多轻量级但设计异构,覆盖了广泛但不均衡的漏洞类型。其有效性主要由模型而非智能体框架决定:Codex/GPT-5.5取得了最大的提升,将检测得分提高了22.8%,捕获奖励提高了43.2%。我们进一步发现,技能触发是一个关键瓶颈。当被触发时,技能保留了共享的六阶段审计工作流程,同时在不同配置下展现出不同的加载模式和对智能体行为的差异化影响。我们发布了我们的技能语料库和工件,以支持未来的研究。

英文摘要

LLM agents, notably Claude Code and OpenAI Codex, are emerging as versatile tools beyond coding agents only. These agents can be enhanced with skills---reusable artifacts that package domain knowledge, workflows, and tool-use instructions. To date, however, little is known about how such skills are designed or how they affect agent effectiveness and behavior in practice. In this paper, we investigate these questions in smart contract security auditing, a domain in which agents have shown substantial promise. We systematically collect 83 smart contract audit skills from the wild and evaluate them on EVMBench across seven agent--model configurations. Our study examines three dimensions: (i) the design characteristics of audit skills, including their structure, knowledge representations, workflows, and tool dependencies; (ii) their effectiveness in improving vulnerability detection; and (iii) their influence on agent execution trajectories. We find that audit skills are mostly lightweight but heterogeneous in design, covering a broad yet imbalanced range of vulnerability types. Their effectiveness is determined primarily by the model rather than the agent harness: Codex/GPT-5.5 achieves the largest gains, improving detection score by 22.8% and captured award by 43.2%. We further find that skill triggering is a key bottleneck. When triggered, skills preserve a shared six-stage audit workflow while exhibiting distinct loading patterns and differential effects on agent behavior across configurations. We release our skill corpus and artifacts to support future research.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑