arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

用于漏洞发现的定向符号执行:KLEE中基于大语言模型的方法

Directed Symbolic Execution for Vulnerability Discovery: An LLM-Guided Approach in KLEE

Lingfeng Chen, Tao Xiao, Masanari Kondo, Yasutaka Kamei

arXiv 2607.21676首次发表:更新:

AI 中文总结

研究针对符号执行的路径爆炸问题,提出基于KLEE的KLEECopilot方法,用大语言模型引导定向符号执行,标记潜在漏洞代码、指导路径优先级并集成循环退出优先级,相比基线提升显著,各组件对有效性有贡献。

AI 中文摘要

符号执行能有效发现安全违规,但存在路径爆炸问题。像KLEE这样的引擎使用路径优先级启发式方法来排序状态探索,通常是优化代码覆盖。然而,路径优先级可能会被困在循环控制流区域中。我们提出了KLEECopilot,一种基于KLEE的大语言模型引导的定向符号执行方法。它用大语言模型标记潜在易受攻击的代码并指导路径优先级,还集成了循环退出优先级以逃离潜在非易受攻击的循环并朝着更深层次的漏洞前进。与Empc等基线相比有显著提升,虽然对模型家族敏感,但对模型规模敏感度低。消融研究表明各组件对有效性有贡献。

英文摘要

Symbolic execution effectively discovers security violations but suffers from path explosion. Engines like KLEE therefore use path prioritization heuristics to order state exploration, typically optimizing code coverage. However, path prioritization can become trapped in cyclic control-flow regions, where repeated branching consumes the exploration budget before exploration reaches vulnerable code beyond these cyclic regions. We propose KLEECopilot, a Large Language Model (LLM)-guided directed symbolic execution approach built on KLEE. KLEECopilot uses LLMs to mark potentially vulnerable code and guide path prioritization. It also integrates loop-exit prioritization to escape potentially non-vulnerable cycles and progress toward deeper vulnerabilities. Compared with baselines such as Empc, KLEECopilot improves basic block coverage by 42.24% and line coverage by 125.82%. It discovers 1,335 total violations and 87 unique violations, outperforming the second-best baseline by 32.2% in total violations and Empc by 24.3% in unique violations. Although KLEECopilot is sensitive to model family, it exhibits only marginal sensitivity to model scale, supporting the efficacy of integrating security semantics and loop-exit prioritization. Ablation studies further show that individual components contribute to effectiveness: alternative configurations involving searchers, internal components, marking sources, and prompt variants yield only 54--61 unique violations, while KLEECopilot maintains competitive code coverage.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑