AI 中文总结
针对企业场景网络钓鱼公开数据集不足的问题,研究团队构建了NEST-Phish合成网络钓鱼邮件数据集,含匹配的合法与钓鱼邮件及标注,可支撑监督检测等多方向研究。
AI 中文摘要
网络钓鱼仍是最持久的网络威胁之一,但针对现实企业邮件场景研究网络钓鱼的公开可共享数据集仍然有限。为解决这一缺口,我们引入围绕虚构组织Nebulon构建的合成企业网络钓鱼邮件数据集。该数据集涵盖广泛的职场沟通主题,包含匹配的合成合法邮件与网络钓鱼邮件,并带有可解释的网络钓鱼线索标注。此处的“合法”指非网络钓鱼类别,而非实际发生的组织邮件。人类受试者分类及分类器评估显示,该数据集支持网络钓鱼判断的有意义变异,同时为监督检测提供可学习信号。这一公开发布的资源旨在支持未来企业类场景下的网络钓鱼检测、人类 susceptibility、可解释性及基准开发相关工作。
英文摘要
Phishing remains one of the most persistent cyber threats, yet publicly shareable datasets for studying phishing in realistic enterprise email settings remain limited. To address this gap, we introduce a synthetic enterprise phishing email dataset built around a fictitious organization, Nebulon. The dataset spans a broad set of workplace communication themes and includes matched synthetic legitimate and phishing emails with interpretable phishing-cue annotations. Here, ``legitimate'' denotes the non-phishing class, not legitimately occurring organizational emails. Human-subject categorizations and classifier evaluations show that the dataset supports meaningful variation in phishing judgments while also providing learnable signal for supervised detection. This publicly released resource is intended to support future work on phishing detection, human susceptibility, explainability, and benchmark development in enterprise-like contexts.
CommentsDataset paper; associated public dataset available on Kaggle: https://www.kaggle.com/datasets/emilywinokur/synthetic-email-corpus