arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.19190cs.CRcs.NIcs.SEcs.SIcs.SYeess.SY

SiNMULI:用于恶意URL识别的新型符号网络方法

SiNMULI: Novel Signed Network Approach for Malicious URL Identification

Avijit Gayen, Sayan Mondal, Angshuman Jana

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出SiNMULI符号网络方法,将恶意URL识别转化为符号网络二分类问题,基于社会平衡理论推理,在真实数据集上以99.89%准确率优于传统模型,具备可解释性等优势,是实用的网络防御方案。

中文摘要 AI 辅助

在人工智能快速发展的当下,计算机安全与在线防护措施已取得显著进步,但恶意网站仍在推动钓鱼计划、欺诈活动及垃圾信息的传播。机器学习、深度学习及仿冒网站检测的传统方法主要依赖静态数据分析,面对不断演变的恶意网络实体时往往效果不佳。针对这些挑战,本研究提出一种基于符号网络的恶意URL识别方法SiNMULI,将有害URL的识别视为根植于社交网络分析基本原理与社会平衡理论的符号网络二分类问题。该方法基于URL的反向链接(即外部超链接)构建符号网络,其中每个节点代表一个URL,超链接作为符号边。利用基于平衡理论的推理机制,本方法通过入链的51%多数规则传播边符号并对未标记域名进行分类。在真实世界数据集上的实验结果显示,SiNMULI的准确率达99.89%、精确率达99.62%、F1值达99.80%,优于传统机器学习与深度学习基准模型。除高准确率外,SiNMULI还具备可解释性、对抗混淆的鲁棒性,且不依赖训练数据,是一种适用于实际网络防御的轻量、可扩展解决方案。

英文摘要

In today's era of rapid advancements in artificial intelligence, computer security and online safeguarding measures have undergone significant improvements. However, malicious websites continue to facilitate the spread of phishing schemes, fraudulent activities and unsolicited communications. Conventional methodologies in machine learning, deep learning and counterfeit website detection predominantly depend on static data analysis, which frequently proves ineffective against the evolving nature of malicious online entities. In response to these challenges, in this work, we propose a signed network-based approach for malicious URL identification, SiNMULI. We introduce an innovative framework that conceptualises the identification of harmful URLs as a signed network-based binary classification problem strongly rooted in the fundamental principles of social network analysis and social balance theory. In this approach, a signed network is constructed based on the backlinks, i.e., external hyperlinks of URLs, wherein each node symbolises a URL and the hyperlinks function as signed edges. Utilising a balance-theoretic inference mechanism, our methodology propagates edge signs and classifies unlabeled domains by employing a 51% majority rule across incoming links. Experimental results on this real-world dataset demonstrate that SiNMULI achieves 99.89% accuracy, 99.62% precision, and 99.80% F1-score, outperforming traditional ML and deep learning baseline models. Beyond high accuracy, SiNMULI offers interpretability, resilience against adversarial obfuscation, and independence from training data, making it a lightweight and scalable solution for real-world cyber defence.

补充信息

↑