arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于匹配器排序:针对威胁报告的知识图谱抽取的可复现性审计

Ranked by the Matcher: A Reproducibility Audit of Knowledge Graph Extraction from Threat Reports

Safayat Bin Hakim, Houbing Herbert Song

arXiv 2609.01671首次发表:更新:

发表机构

University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对威胁报告知识图谱抽取工具的可复现性展开审计,构建CTIForge分离组件效应,发现匹配器会影响F1分数排序,LLM评审器一致性优于机械匹配器,相关成果已公开。

AI 中文摘要

安全团队和研究人员选择用于威胁报告的知识图谱抽取工具时,会参考已公布的三元组F1分数,但这些分数取决于预测的三元组与黄金标注的匹配方式。在被检查的12个系统中,我们仅能重新实现其中5个系统所声明的匹配规则。在8种协议下对10个系统在共享文档上的输出重新评分,逆转了45组成对排序中的11组;某一固定预测集的F1值范围为0.16至0.70。在GRID的外部378项校准集上,无机械匹配器(词汇型、嵌入型或蕴含型)与多评审员裁决的一致性超过71%,而大型语言模型(LLM)评审器达到86%。为分离组件效应与匹配器奖励,我们构建了CTIForge,其确定性验证层可在抽取结果保持字节级相同的情况下发生变化。在7种测试部署配置中,验证层提升了全部4个托管 backbone 的精确率,同时降低了全部3个离线 backbone 的精确率。由于 backbone、解码及后端特定提示存在共变,这是一种描述性划分而非孤立服务效应。这与明确质疑实体类型的操作约2.8倍的增长相吻合,符合编码了其开发所针对的抽取器约定的手写规则。我们发布了该流水线、协议套件及逐三元组审计记录。

英文摘要

Security teams and researchers choose knowledge-graph extraction tooling for threat reports on the strength of published triple-F1 scores, yet those scores depend on how predicted triples are matched to gold annotations. We could reimplement the stated matching rule for only five of twelve inspected systems. Re-scoring ten system outputs on shared documents under eight protocols reverses eleven of forty-five pairwise orderings; one fixed prediction set spans 0.16--0.70 F1. On an external, human-adjudicated set, no mechanical matcher---lexical, embedding, or entailment---agrees with the reviewers more than seven times in ten; an LLM judge agrees far more often. To separate component effects from matcher rewards, we build CTIForge, whose deterministic validation layer can vary while extraction is held byte-identical. Across seven tested deployment configurations, no hosted backbone loses precision under validation and every offline one does. Backbone, decoding, and prompting covary across those configurations; on the one backbone we could serve both ways, serving alone reproduces the split. It coincides with a roughly 2.8-fold increase in actions explicitly disputing entity type, consistent with hand-written rules encoding the conventions of the extractor against which they were developed. We release the pipeline, protocol suite, and per-triple audit records.

CommentsCode and configurations are available at https://github.com/sbhakim/CTIForge

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑