arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.16158cs.LGcs.AIcs.ET

一种针对钓鱼模板与攻击者归因分析的树结构方法

A Tree-Structured Approach for Phishing Template and Attacker Attribution Analysis

Unai Agirre, Imanol Jerico, Felipe Castaño, Andrea Venturi, Francesco Zola

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出一种基于DOM树结构的无监督聚类方法,用于识别钓鱼模板复用,可检测新兴及零日钓鱼模板并分析协同钓鱼威胁。

中文摘要 AI 辅助

钓鱼仍是持续且不断演变的网络安全威胁,攻击量已达到创纪录水平。这种增长由钓鱼的产业化驱动,得益于广泛可用的钓鱼工具包(phishing kits)和可重复使用的模板,使网络犯罪分子能够快速生成并部署大量欺诈性网页。尽管这些网站的表层属性可能存在差异,但它们的底层结构通常表现出显著相似性。然而,大多数现有防御措施依赖于反应式黑名单或监督分类模型,这些模型聚焦于单个钓鱼实例,限制了其识别结构复用和检测协同钓鱼活动的能力。为解决这一局限,本研究探究HTML结构是否可作为识别钓鱼模板复用的鲁棒指纹。我们将网页建模为文档对象模型(DOM)树,提取结构特征,可选地补充基于HTML标签的内容信息。随后使用无监督学习方法对这些表示进行聚类,以将结构相似的网页分组。我们评估并比较了三种聚类算法,同时分析提取的DOM树深度如何影响聚类形成和整体聚类性能。最后,还通过定量和定性方式评估聚类质量,包括一种新颖的分层Jaccard距离分数以及可视化工具支持的人工检查。结果表明,网页的结构表示可有效揭示钓鱼网站之间隐藏的相似性,从而能够检测新兴和零日模板,并支持对协同钓鱼威胁的分析。

英文摘要

Phishing remains a persistent and evolving cybersecurity threat, with attack volumes reaching record levels. This growth is driven by the industrialization of phishing through widely available phishing kits and reusable templates, which enable cybercriminals to rapidly generate and deploy large numbers of fraudulent webpages. Although surface-level attributes may differ across these websites, their underlying structures often exhibit significant similarities. However, most existing defenses rely on reactive blocklists or supervised classification models that focus on individual phishing instances, limiting their ability to identify structural reuse and detect coordinated phishing campaigns. To address this limitation, this study investigates whether HTML structure can serve as a robust fingerprint for identifying phishing template reuse. We model webpages as Document Object Model (DOM) trees and extract structural features, optionally enriched with HTML tag-based content information. These representations are then clustered using unsupervised learning methods to group structurally similar webpages. Three clustering algorithms are evaluated and compared, while also analyzing how the depth of the extracted DOM-tree affects cluster formation and overall clustering performance. Finally, cluster quality is also evaluated both quantitatively and qualitatively, including a novel level-wise Jaccard Distance Score and manual inspection supported by visualization tools. Results demonstrate that structural representations of webpages can effectively reveal hidden similarities across phishing sites, enabling the detection of emerging and zero-day templates and supporting the analysis of coordinated phishing threats

发表机构

  • Vicomtech (BRTA)(维科姆科技(BRTA))
  • University of León(莱昂大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑