Token去向何处?面向漏洞发现的大语言模型智能体的成本理解与降低
Where Do the Tokens Go? Understanding and Reducing Costs in LLM Agents for Vulnerability Discovery
浏览论文内容
中文总结 AI 辅助
该研究针对LLM智能体漏洞发现中Token消耗高却难出成果的问题,通过分析200条CyberGym轨迹揭示瓶颈,提出AVRI接口,可降低Codex和OpenCode的Token消耗并保持性能。
中文摘要 AI 辅助
大语言模型(LLM)智能体在漏洞发现过程中可能消耗数百万个Token,却无法生成可工作的概念验证(PoC)。这些预算被什么消耗了?为何无法产生结果?我们通过对200条CyberGym轨迹的多轴开放编码研究,诊断这些成本与失败原因,研究覆盖四个智能体(即Codex、OpenCode、Cybench和EnIGMA),包含无辅助基线及四种现有效率方法。研究揭示三个关键发现:第一,不同智能体的成功率和成本差异显著,更高的Token消耗并不总能带来更好结果;第二,代码定位与理解,结合漏洞推理与触发设计,占Token消耗的60.4%,是失败运行的两大主要瓶颈;第三,仅24.4%的匹配比较能在更低总Token消耗下保持成功率,不合适的信号与辅助开销限制了现有方法的效益。基于这些发现,我们提出AVRI,一种以智能体为中心的漏洞推理接口,围绕持久双向证据轨迹(BET)构建。BET连接测试框架如何消耗输入与使选定操作不安全所需的条件,保留源支持的对应关系,以及智能体的假设与未解决问题。读取、分析和持久化命令帮助智能体构建并复用该证据,而非反复检索与重建。在20个评估任务上,AVRI使Codex的总Token消耗降低18.0%,OpenCode降低23.7%,同时保持成功率并提升或维持召回率。
英文摘要
LLM agents can spend millions of tokens during vulnerability discovery without producing a working proof of concept (PoC). What consumes that budget, and why does it fail to produce results? We diagnose these costs and failures through a multi-axis open-coding study of 200 CyberGym traces, spanning four agents (i.e., Codex, OpenCode, Cybench, and EnIGMA) under an unaided baseline and four existing efficiency methods. The study reveals three key findings. First, different agents vary substantially in success and cost, and higher spending does not consistently yield better outcomes. Second, code localization and understanding, together with vulnerability reasoning and trigger design, account for 60.4% of tokens and represent the two leading bottlenecks in failed runs. Third, only 24.4% of matched comparisons preserve success at lower total cost; unsuitable signals and auxiliary overhead limit the benefits of existing methods. Motivated by these findings, we present AVRI, an Agent-centric Vulnerability Reasoning Interface built around a persistent Bidirectional Evidence Trace (BET). BET connects how the harness consumes input with the conditions required to make a selected operation unsafe, retaining source-supported correspondences alongside the agent's hypotheses and open questions. Reading, analysis, and persistence commands help agents build and reuse this evidence rather than repeatedly retrieve and reconstruct it. On 20 evaluation tasks, AVRI reduces total cost by 18.0% for Codex and 23.7% for OpenCode while preserving success rates and improving or maintaining recall.
发表机构
- Huazhong University of Science and Technology(华中科技大学)
机构由 AI 辅助整理,请以论文原文为准。