arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PhiShark2026:用于钓鱼网站研究的多层主动网络原始证据数据集

PhiShark2026: A Multi-Layer Active-Web Raw-Evidence Dataset for Phishing Website Research

Furkan Çolhak, Ferhat Demirkıran, Hasan Dağ, Alexander Iliev

arXiv 2608.23199首次发表:更新:

AI 中文总结

本研究构建PhiShark2026多层主动网络原始证据数据集,含67502次扫描数据,采用托管感知证据模型处理共享平台问题,揭示钓鱼与良性网站的系统性差异,为钓鱼研究提供可检查可复现的基础。

AI 中文摘要

钓鱼网站生命周期短且变化迅速,但许多钓鱼网站数据集将观测结果简化为URL或预计算特征,限制研究者使用预定义表示,丢弃了推导替代特征、应用新提取方法、检查跨层关系及随钓鱼技术演变重新分析观测所需的底层证据。本研究解决该局限,构建含67502次扫描的多层主动网络数据集,其中包括来自运营数据源的33387个钓鱼观测结果和34115个经筛选的良性参考观测结果。语料库保留HTML内容、截图、URL与重定向行为、HTTP及安全头、合规文件、TLS证书、DNS与域名注册、开放端口、地理位置与可访问性测量、网络基础设施等各层原始证据,且明确记录不可用证据而非将其视为负观测。为避免共享平台上基础设施归因误导,本研究采用托管感知证据模型,对免费托管租户页面屏蔽提供商所有的基础设施信号,同时保留有意义的页面级和传输级证据。特征分析显示,钓鱼网站与良性网站在网络资源使用、域名成熟度、邮件与策略配置、安全头及基础设施上下文方面存在系统性差异。通过保留原始人工制品、采集元数据及明确的证据可用性,该语料库为未来钓鱼测量和数据集研究提供了可检查、可复现的基础。

英文摘要

Phishing websites are short-lived and rapidly changing, yet many phishing datasets reduce observations to URLs or precomputed features, constraining researchers to predefined representations and discarding the underlying evidence needed to derive alternative features, apply new extraction methods, examine cross-layer relationships, and reanalyze observations as phishing techniques evolve. This study addresses this limitation with a multi-layer active-web dataset comprising 67,502 scans, including 33,387 phishing observations from operational feeds and 34,115 screened benign reference observations. The corpus preserves raw evidence across HTML content and screenshots, URL and redirect behavior, HTTP and security headers, compliance files, TLS certificates, DNS and domain registration, open ports, geolocation and accessibility measurements, and network infrastructure, while explicitly recording unavailable evidence rather than treating it as negative observations. To avoid misleading infrastructure attribution on shared platforms, the study applies a hosting-aware evidence model that masks provider-owned infrastructure signals for free-hosted tenant pages while retaining meaningful page- and transport-level evidence. Characterization reveals systematic differences between phishing and benign websites across web-resource usage, domain maturity, mail and policy configuration, security headers, and infrastructure context. By preserving raw artifacts together with acquisition metadata and explicit evidence availability, the corpus provides an inspectable and reproducible foundation for future phishing measurement and dataset research.

Comments13 pages, 4 figures, 13 tables

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑