arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

经典模型与基于Transformer的模型在文档敏感性分类任务中的基准测试

Benchmarking Classical and Transformer-Based Models for Document Sensitivity Classification

Aleesha Zainab, Muhammad Ahmed Khalid, Faheem Ullah Khan, Asifullah Khan

arXiv 2608.16928首次发表:更新:

AI 中文总结

该研究针对文档敏感性分类的标签泄漏问题,构建了含16000份外交电报的Strategic 16K基准语料库,测试发现BERT性能最优、TF-IDF结合逻辑回归性价比最高,成果为该领域提供了首个可复现的泄漏控制基准。

AI 中文摘要

组织文档的自动敏感性分类是一个关键却未得到充分关注的问题,分类错误的后果从违规到安全漏洞不等。基于AI的方法为人工审核提供了可扩展的替代方案,但其可靠性根本上取决于训练数据的完整性。该领域存在一个普遍但未被充分报道的问题:标签泄漏,即文档主体中嵌入的残留分类标记,使模型能够利用表面捷径而非学习真正基于内容的敏感性信号,从而产生被夸大且不可靠的性能估计。本文通过引入Strategic 16K解决了这一问题,这是一个精心构建的、经泄漏控制的语料库,包含来自维基解密美国外交公共图书馆(PlusD)的16000份外交电报,并提供了系统的基准测试,评估了涵盖经典机器学习和基于Transformer方法的六种模型架构。我们记录了一个扩展的泄漏去除协议,该协议识别并消除了文档主体中嵌入的三类残留分类标记。在干净的基准测试中,BERT实现了最强性能(准确率=89.14%,F1值=89.33%),其次是ELECTRA(准确率=88.57%,F1值=88.90%)。在经典模型中,结合逻辑回归的TF-IDF实现了最强性能,且计算成本显著更低。这些结果构成了首个在明确的泄漏控制条件下,从维基解密PlusD构建的完全可复现的敏感性分类基准。

英文摘要

Automatic sensitivity classification of organizational documents is a critical yet underserved problem, where the consequences of misclassification range from regulatory violations to security breaches. While AI-based approaches offer a scalable alternative to manual review, their reliability depends fundamentally on the integrity of training data. A pervasive but underreported problem in this domain is label leakage: residual classification markers embedded within document bodies that allow models to exploit surface shortcuts rather than learning genuine content-based sensitivity signals, producing performance estimates that are inflated and unreliable. This paper addresses this problem by introducing Strategic 16K, a carefully constructed, leakage-controlled corpus of 16,000 diplomatic cables sourced from the WikiLeaks Public Library of US Diplomacy (PlusD), and presents a systematic benchmark evaluating six model architectures spanning classical machine learning and transformer-based approaches. We document an extended leakage removal protocol that identifies and eliminates three categories of residual classification markers embedded within document bodies. On the clean benchmark, BERT achieves the strongest performance (Accuracy = 89.14%, F1 = 89.33%), followed by ELECTRA (Accuracy = 88.57%, F1 = 88.90%). Among classical models, TF-IDF with Logistic Regression achieves the strongest performance at significantly lower computational cost. These results constitute the first fully reproducible sensitivity classification benchmark constructed under explicit leakage-controlled conditions from WikiLeaks PlusD.

Comments6 pages, 4 images

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑