曝光度最高的 Docker Hub 镜像中的漏洞、秘密与配置错误
Vulnerabilities, Secrets and Misconfiguration in the Highest-Exposure Docker Hub Images
浏览论文内容
中文总结 AI 辅助
本文提出 ChimangoScan 管道,爬取 Docker Hub 高曝光镜像并使用多扫描器分析,发现高曝光镜像普遍存在漏洞、配置错误及误报的秘密,发布了相关管道与数据集。
中文摘要 AI 辅助
Docker Hub 是大多数容器部署底层的镜像仓库,广泛复用的基础镜像存在缺陷时,所有基于该基础镜像构建的镜像都会继承这一缺陷。此前针对整个生态系统的测量均依赖单一检测器,导致其计数的工具依赖性未被量化;而对比扫描器的研究仅使用数十到数百张镜像的样本。本文提出 ChimangoScan,这一管道可爬取 Docker Hub 命名空间(共 12,716,568 个仓库,累计拉取量达 6638 亿次),重构镜像层图(含 5440 万条 IS_BASE_OF 边),并将镜像自身拉取量及其整个下游子树的拉取量整合为一个标量,以此计算曝光度得分对镜像进行排名;随后使用 6 种独立扫描器对曝光度最高的 52,895 个仓库(占所有已记录拉取量的 84.7%)进行扫描,共产生 1.704 亿条发现结果。漏洞几乎普遍存在:96.3% 的镜像带有已知软件包漏洞,93.4% 的镜像带有严重漏洞,98.0% 的镜像至少存在一项 CIS Docker Benchmark 配置错误。单一工具报告的安全态势在很大程度上是该工具的产物:在 8070 万个不同的(漏洞、软件包)组中,66.8% 仅被三个漏洞扫描器中的一个标记,仅 2.7% 被所有三个扫描器标记,最佳单一扫描器的覆盖率为 66.9%。TruffleHog 在 76.9% 的镜像中标记出秘密,但对 1100 个随机检测结果的手动标注发现,99.7% 并非凭证。一个单一的 zlib CVE 影响了占整个数据集曝光度 47.3% 的镜像,并传播到 113 万个不同的下游镜像,但曝光度无法预测镜像的脆弱程度。本文发布了该管道及 283 GB 的数据集。
英文摘要
Docker Hub is the registry underneath most container deployments, and a flaw in a widely reused base image is inherited by every image built on it. Prior ecosystem-scale measurements each rely on a single detector, leaving the tool-dependence of their counts unquantified, while the studies that do compare scanners use samples of tens to hundreds of images. We present ChimangoScan, a pipeline that crawls the Docker Hub namespace (12,716,568 repositories, 663.8 billion cumulative pulls), reconstructs the image layer graph (54.4 million IS_BASE_OF edges), ranks images by an exposure score that folds an image's own pull count and those of its entire downstream subtree into one scalar, and scans the 52,895 highest-exposure repositories (84.7% of all recorded pulls) with six independent scanners, yielding 170.4 million findings. Vulnerabilities are near-universal: 96.3% of images carry a known package vulnerability, 93.4% a critical one, and 98.0% at least one CIS Docker Benchmark misconfiguration. The posture a single tool reports is largely an artifact of that tool: of 80.7 million distinct (vulnerability, package) groups, 66.8% are flagged by only one of the three vulnerability scanners and just 2.7% by all three, and the best single scanner recovers 66.9%. TruffleHog flags a secret in 76.9% of images, yet hand-labeling 1,100 random detections finds 99.7% are non-credentials. A single zlib CVE reaches images carrying 47.3% of total corpus exposure and propagates to 1.13 million distinct downstream images, but exposure does not predict how vulnerable an image is. We release the pipeline and the 283 GB dataset.
发表机构
- AI Horizon Labs(AI地平线实验室)
- Federal University of Pampa (UNIPAMPA)(潘帕联邦大学)
机构由 AI 辅助整理,请以论文原文为准。