发表机构
OWASP GenAI Security Project — Top 10 for LLM Applications(OWASP GenAI安全项目——LLM应用十大项目)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究以OWASP 2026年LLM应用Top 10专家排名为对象,结合7714条LLM安全事件语料库,用贝叶斯模型分析其鲁棒性,发现两者一致性弱但专家排名仍具鲁棒性,且未推翻官方共识。
AI 中文摘要
OWASP Top 10 for LLM应用是由安全从业者社区评定的最重要风险排名。本文提出更具体的问题:对照真实事件记录,该专家排名是否与数据相符?我们整理了来自CVE、GHSA、OSV和AIAAIC的大规模LLM安全事件语料库,共7714条快照数据,其中6639条按20项分类法标注,并使用贝叶斯测量误差模型推导基于事件的排名,该模型会针对分类器的精确率和召回率修正每个类别的计数。2026年候选名单以固定权重融合两种信号,专家投票占0.75,数据占0.25,因此语料库在不推翻共识的前提下修正了共识。两种排名的一致性较弱:Cohen's κ≈0.20,90%置信区间包含零。不过专家排名具有鲁棒性:对四个前沿分类器的预注册对比未得出胜者,均未超过事件基准的平衡准确率0.863;真实值验证显示,基准的排序(与保留真实值的Spearman ρ=0.918)保持有效。需说明,这是两名工作组成员的探索性分析,非OWASP官方发布,不取代官方列表或流程。
英文摘要
The OWASP Top 10 for LLM Applications ranks the risks that a community of security practitioners judges most important. We ask a narrower question: checked against the record of real incidents, does that expert ranking agree with the data? We assembled a large-scale corpus of LLM-security incidents (7,714 snapshotted and 6,639 labeled against the 20-entry taxonomy) drawn from CVE, GHSA, OSV, and AIAAIC, and derived an incident-based ranking with a Bayesian measurement-error model that corrects each category's count for classifier precision and recall. The 2026 candidate list blends the two signals at fixed weights, 0.75 on the expert vote and 0.25 on the data, so the corpus corrects the consensus without overturning it. The agreement between the two rankings is weak: Cohen's $κ\approx 0.20$, with a 90% interval that crosses zero. The expert ranking is nonetheless robust. A pre-registered bake-off of four frontier classifiers returns no winner. None beats the incidence floor's balanced accuracy of 0.863. A ground-truth check leaves the floor's ordering (Spearman $ρ= 0.918$ against held-out truth) in place. This is an exploratory analysis by two working-group members, not the official OWASP release, and it does not supersede the official list or process.
Comments26 pages, 19 figures. Exploratory incident-data analysis by two members of the OWASP GenAI Security Project Top 10 for LLM Applications working group. Analysis predates the official OWASP GenAI LLM Top 10 2026 (published August 2026). Not an official OWASP release and does not supersede the official list or process. Code and data: https://github.com/rocklambros/incident-rank-validation