发表机构
San Diego State University(圣地亚哥州立大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出一种基于YOLOv8的视觉钓鱼检测方法,通过分析网页截图的外观特征进行分类,达到92%准确率,处理速度约100毫秒,数据增强后假阳性率降低11%。
AI 中文摘要
钓鱼攻击仍然是金融和身份欺诈最常见的载体之一,而大多数检测系统仍依赖于检查页面的URL、HTML标记或域名注册历史。这些信号容易被攻击者轮换或混淆,而且它们几乎无法说明真正说服受害者交出密码或卡号的因素:页面的外观。本文描述了一种基于视觉、图像处理的钓鱼检测方法,将渲染后的网页视为人类眼睛所看到的一幅图片,要么匹配可信品牌,要么不匹配。训练了一个YOLOv8卷积神经网络,根据布局、标志放置、配色方案和登录表单结构,而非从页面提取的文本,将整页网站截图分类为钓鱼或合法。该系统在保留测试集上达到了92%的分类准确率,处理单个截图大约需要100毫秒,并且在进行了一轮专门针对光照、压缩和缩放变化的数据增强后,相对于增强前的基线,将假阳性率降低了11%。本文详细介绍了数据集构建、增强策略、模型架构和训练设置以及最终性能,并最后讨论了这种视觉检测器如何与现有的基于URL和内容的防御措施相辅相成,而非替代它们。
英文摘要
Phishing remains one of the most common vectors for financial and identity fraud, and most detection systems still rely on inspecting a page's URL, HTML markup, or domain registration history. These signals are easy for an attacker to rotate or obfuscate, and they say very little about what actually convinces a victim to hand over a password or a card number: the way the page looks. This paper describes a visual, image-based approach to phishing detection that treats a rendered webpage the same way a human eye would, as a picture that either matches a trusted brand or doesn't. A YOLOv8 convolutional neural network was trained to classify full-page website screenshots as phishing or legitimate based on layout, logo placement, color scheme, and login-form structure, rather than on text extracted from the page. The system reached 92% classification accuracy on a held-out test set, processed a single screenshot in roughly 100 milliseconds, and, after a round of data augmentation aimed specifically at lighting, compression, and scaling variation, cut the false-positive rate by 11% relative to the pre-augmentation baseline. The paper walks through the dataset construction, the augmentation strategy, the model architecture and training setup, and the resulting performance, and closes with a discussion of where this kind of visual detector fits alongside, rather than instead of, existing URL- and content-based defenses.