发表机构
Carnegie Mellon University(卡内基梅隆大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM漏洞分析中自由形式推理易掩盖错误的问题,提出VERA框架,通过结构化推理记录和确定性检查审计推理,能暴露87%被传统LLM评判遗漏的错误。
AI 中文摘要
大型语言模型(LLMs)正越来越多地被部署用于自动化软件漏洞分析。仅靠二元分类是不够的;实践者需要解释来对漏洞进行分类(triage)并设计补丁。标准实践依赖于思维链(Chain-of-Thought, CoT)提示,但自由形式的推理允许模型用看似合理的叙述掩盖逻辑跳跃、幻觉执行步骤和内部不一致。我们的手动审计揭示,大约60%的正确漏洞判决伴随着捏造或无法验证的声明,且自由形式的解释使推理错误能够逃避LLM作为评判者(LLM-as-a-judge)的评估。我们提出了漏洞解释推理审计器(Vulnerability Explanation Reasoning Auditor, VERA),一个用于审计LLM漏洞推理的自动化框架。VERA不接受自由形式的文本,而是要求模型输出一个结构化推理记录(Structured Reasoning Record, SRR),该记录以机器可读字段编码跟踪指针、内存操作和状态转换。一个多阶段评判者使用确定性检查针对八种推理失败模式审计每个SRR,LLM调用仅保留用于语义解释。标准化的SRR模式还支持自动化变异测试,以便在没有人工标注的情况下大规模基准测试评判者。我们的评估显示,推理缺陷在正确判决中的出现频率与错误判决中一样高,且VERA暴露了87%的自由形式LLM评判者系统性地遗漏的推理错误。
英文摘要
Large Language Models (LLMs) are increasingly deployed for automated software vulnerability analysis. Binary classification alone is insufficient; practitioners need explanations to triage bugs and engineer patches. Standard practice relies on Chain-of-Thought (CoT) prompting, but free-form reasoning allows models to obscure logical leaps, hallucinated execution steps, and internal inconsistencies behind plausible prose. Our manual audit reveals that approximately 60% of correct vulnerability verdicts are accompanied by fabricated or unverifiable claims, and free-form explanations allow reasoning errors to evade LLM-as-a-judge evaluation. We present Vulnerability Explanation Reasoning Auditor (VERA), an automated framework for auditing LLM vulnerability reasoning. Rather than accepting free-form text, VERA asks models to output a Structured Reasoning Record (SRR) encoding tracked pointers, memory operations, and state transitions in machine-readable fields. A multi-stage judge audits each SRR against eight reasoning failure modes using deterministic checks, with LLM calls reserved for semantic interpretation. The standardized SRR schema also enables automated mutation testing to benchmark judges at scale without human annotation. Our evaluation shows reasoning flaws occur in correct verdicts just as frequently as incorrect ones, and VERA exposes 87% of reasoning errors that free-form LLM-as-judge systematically miss.