arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.10735cs.CRcs.SE

DITTO:一种用于有效安全审计的基于Pickle的上下文感知预训练模型扫描器

DITTO: A Context-aware Pickle-based Pre-Trained Model Scanner for Effective Security Audits

Qiaolin Qin, Wanpeng Li, Benoit Baudry, Lorenzo De Carli, Heng Li, Ettore Merlo

首次发表
浏览论文内容

中文总结 AI 辅助

DITTO是首个基于栈的上下文感知Pickle预训练模型扫描器,通过跟踪虚拟机状态转换与语义分析,在PickleBench基准上实现100%覆盖、0漏报,显著提升安全审计性能。

中文摘要 AI 辅助

预训练模型(PTMs)常以序列化二进制文件形式分发,但其复用常使软件供应链面临反序列化攻击风险。尽管更安全的序列化格式已出现,不安全的Pickle格式仍普遍存在:我们对超10000个热门Hugging Face仓库的分析显示,9.3%的仓库依赖Pickle。虽已提出多种防御机制,但最先进的模型扫描器存在覆盖度-精度差距,遗漏安全敏感行为并产生过多误报。本文提出DITTO,首个基于栈的、针对基于Pickle的PTMs的上下文感知扫描器。DITTO忠实跟踪Pickle虚拟机状态转换,执行上下文感知语义分析以推断模型意图。我们还提出PickleBench,包含959个良性和92个真实恶意模型的基准,涵盖现有工具此前遗漏的扩展注册表攻击。多项评估显示,DITTO实现100%扫描覆盖度、0%漏报率、0.7%误报率,F1分数达0.966,显著优于最先进扫描器。通过最小化误报同时保持检测精度,DITTO生成带上下文证据的可操作安全报告,支持安全的PTMs复用并增强软件供应链完整性。

英文摘要

Pre-trained models (PTMs) are widely distributed as serialized binaries, but their reuse often exposes software supply chains to deserialization attacks. Despite the emergence of safer serialization formats, the unsafe Pickle format remains prevalent: our analysis of over 10,000 popular Hugging Face repositories reveals that 9.3% rely on Pickle. While many defense mechanisms have been proposed, state-of-the-art model scanners suffer from a coverage-precision gap, missing security-sensitive behaviors and generating excessive false alerts. In this paper, we introduce DITTO, the first stack-based, context-aware scanner for Pickle-based PTMs. DITTO faithfully tracks Pickle virtual machine state transitions and performs context-aware semantic analysis to infer model intentions. We also present PickleBench, a benchmark of 959 benign and 92 malicious real-world models, including extension registry attacks previously missed by existing tools. Across multiple evaluations, DITTO achieves 100% scanning coverage, a 0% false-negative rate, and a 0.7% false-positive rate, yielding an F1 score of 0.966, significantly outperforming state-of-the-art scanners. By minimizing false alerts while preserving detection accuracy, DITTO generates actionable security reports with contextual evidence, enabling safe PTM reuse and strengthening software supply chain integrity.

发表机构

  • Polytechnique Montreal(蒙特利尔理工学院)
  • University of Liverpool(利物浦大学)
  • Université de Montréal(蒙特利尔大学)
  • University of Calgary(卡尔加里大学)

机构由 AI 辅助整理,请以论文原文为准。

↑