arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01758cs.CRcs.SE

基于行为树引导的轻量级大语言模型漏洞检测研究

Towards Behavior Tree-Guided Vulnerability Detection with Lightweight LLMs

Enna Basic, Alberto Giaretta

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出用行为树(BT)作为中间表示,结合量化本地LLM Mistral Small 3.2 24B(Q4_K_M),在460个Java样本上验证其可提升漏洞检测性能,适配上下文窗口,为漏洞检测提供紧凑结构化表示。

中文摘要 AI 辅助

大语言模型(LLMs)越来越多地被用于软件漏洞检测,但其性能取决于源代码在输入中的表示方式。大多数提示方法使用原始形式的源代码,而一些研究提出使用结构化表示。抽象语法树(ASTs)是最流行的方法之一,但AST的冗余度相对于源代码增加了输入大小,使其难以适配某些LLMs的上下文窗口。本文研究行为树(BTs)作为基于LLM的漏洞检测的替代中间表示。BTs比ASTs更紧凑地编码控制流、条件和可执行动作,在标记数受限的情况下是自然的候选方案。首先,我们提出一个预处理阶段,将Java源代码解析为ASTs,然后将其转换为BT表示。接着,我们使用三种输入表示(原始源代码、AST和BT),比较来自Juliet Java测试套件的460个Java样本的漏洞检测性能。所有实验均使用单个量化本地LLM——Mistral Small 3.2 24B(Q4_K_M)。我们的结果显示,使用BT表示可提高短代码样本的召回率,而原始源代码实现更高的精度。在较长样本上,BT比原始表示提升了整体性能且适配上下文窗口,而许多AST超出了上下文限制。这些发现表明,BT可为使用可量化、可本地部署的LLM进行漏洞检测提供紧凑且有用的结构化表示。

英文摘要

Large Language Models (LLMs) are increasingly used for software vulnerability detection, but their performance depends on how source code is represented in the input. Most prompting approaches use source code in its original form, while some works propose the use of structured representations. Abstract Syntax Trees (ASTs) are one of the most popular approaches, but AST verbosity increases input size relative to source code, making them hard to fit within some LLMs context windows. This paper investigates Behavior Trees (BTs) as an alternative intermediate representation for LLM-based vulnerability detection. BTs encode control flow, conditions, and executable actions more compactly than ASTs, making them a natural candidate when token count is a constraint. First, we propose a preprocessing stage that parses Java source code into ASTs and then converts them into BT representations. We then compare vulnerability detection performance across 460 Java samples from the Juliet Java test suite, using three input representations: raw source code, AST, and BT. All experiments use a single quantized local LLM, Mistral Small 3.2 24B (Q4_K_M). Our results show that using BT representations improves recall on short code samples, while raw source code achieves higher precision. On longer samples, BTs improve overall performance over the original representation and fit within the context window, whereas many ASTs exceed the context limit. These findings suggest that BTs can provide a compact and useful structured representation for vulnerability detection with quantized, locally deployable LLMs.

发表机构

  • Örebro University(欧雷布罗大学)

机构由 AI 辅助整理,请以论文原文为准。

↑