超越准确率:面向结构化虚假招聘信息检测的质心引导对比损失
Beyond Accuracy: Centroid-Guided Contrastive Loss for Structured Fraudulent Job Posting Detection
浏览论文内容
中文总结 AI 辅助
针对虚假招聘信息检测,提出质心引导对比损失(CGCL),统一分类与聚类,通过质心驱动的top-k推拉机制重塑潜在空间,在EMSCAD基准上取得SOTA性能。
中文摘要 AI 辅助
虚假招聘信息检测旨在识别因虚假内容、误导性信息或负面意图而受损的招聘广告,这些广告扰乱了求职者和雇主的在线生态系统。该领域的现有研究缺乏有效方法,无法同时实现高准确率和具有意义的潜在空间表征结构,以捕捉虚假帖子之间的细微差别。为此,我们提出了质心引导对比损失(CGCL),这是一种损失函数,它将分类与密集形式的聚类统一起来,通过质心驱动的top-$k$推拉机制持续重塑潜在空间。CGCL的互补特性使模型能够强制执行准确的决策边界并保持高聚类紧凑性,有效捕捉类别可分性和潜在结构。大量实验表明,我们的方法在公共基准数据集EMSCAD上达到了最先进的(SOTA)性能。与本研究相关的代码可在以下网址获取:此https URL
英文摘要
Fraudulent job posting detection aims to identify job advertisements that are corrupted either through fake content, misleading information, or negative intent, disrupting the online eco-system of job-seekers and employers. Existing studies in this domain lack effective methods to simultaneously achieve high accuracy and meaningful structure of latent-space representations that capture subtleties among fake posts. To this end, we propose Centroid-Guided Contrastive Loss (CGCL), a loss function which unifies classification with densely formulated clustering to consistently reshape latent-space through a centroid-driven top-$k$ push-and-pull mechanism. The complementary nature of CGCL enables the model to enforce accurate decision boundaries and maintain high clustering compactness, effectively capturing both class separability and latent structure. Extensive experiments demonstrate the state-of-the-art (SOTA) performance of our method on EMSCAD, a public benchmark dataset. The code associated with this work is available at: https://github.com/ali-ahmed925/CGCL_code
发表机构
- National University of Computer and Emerging Sciences(国立计算机与新兴科学大学)
- Islamic University of Madinah(麦地那伊斯兰大学)
机构由 AI 辅助整理,请以论文原文为准。