arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

检测欺骗性招聘:基于信号理论的机器学习框架用于早期识别劳动剥削

Detecting Deceptive Recruitment: A Signal-theoretic Machine Learning Framework for Early Identification of Labour Exploitation

Sajid Siraj, Mahnaz Hosseinzadeh, Amin Vafadarnikjoo, Shuyang Li

arXiv 2609.20336首次发表:更新:

发表机构

University of Leeds; COMSATS University; University of Sheffield; University of Birmingham(利兹大学; COMSATS大学; 谢菲尔德大学; 伯明翰大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出基于信号理论的机器学习框架,利用多模态特征早期识别欺骗性招聘广告,以应对劳动剥削,通过验证案例和特征分析实现高判别性能。

AI 中文摘要

欺骗性的在线招聘广告已成为进入强迫劳动的主要途径,然而由于数据稀缺和缺乏经验验证的指标,系统性检测方法仍不完善。我们将这一检测挑战形式化为信号理论下的分类问题,其中剥削者传输模仿合法通信的无成本信号,涉及文本、视觉和结构维度。利用通过反奴隶制慈善机构在九个来源国和21个行业中收集的464个验证案例(164个欺骗性,300个合法),我们开发了结合计算机视觉、自然语言处理和语义嵌入的多模态检测模型。通过系统性特征消融实验和重复分层交叉验证,我们证明单一模态具有显著的判别能力(ROC-AUC:0.87--0.97),而它们的整合带来了适度的进一步提升。基于SHAP的分析揭示,文本质量和领域特定风险语言是主要判别因素,其中可读性指数、风险关键词密度和签证赞助提及排名最高,其次是视觉颜色和纹理特征。这些生产质量差距反映了资源限制,使得剥削者无法在所有通信渠道上同时维持专业标准。我们通过一个概念验证决策支持系统将研究结果操作化,为从业者提供可解释的风险评分。这项工作展示了严谨的分析框架如何应对以信息不对称和有限地面真实数据为特征的复杂人道主义行动挑战。

英文摘要

Deceptive online job advertisements have emerged as a primary pathway into forced labour, yet systematic detection methods remain underdeveloped due to data scarcity and absence of empirically validated indicators. We formalise this detection challenge as a classification problem under signalling theory, where exploiters transmit costless signals mimicking legitimate communications across textual, visual, and structural dimensions. Using 464 verified cases (164 deceptive, 300 legitimate) collected through anti-slavery charities across nine origin countries and 21 industries, we develop multimodal detection models combining computer vision, natural language processing, and semantic embeddings. Through systematic feature ablation experiments and repeated stratified cross-validation, we demonstrate that individual modalities achieve substantial discriminatory power (ROC-AUC: 0.87--0.97), whilst their integration yields modest further gains. SHAP-based analysis reveals that text quality and domain-specific risk language are the primary discriminators, with readability indices, risk keyword density, and visa sponsorship mentions ranking highest, followed by visual colour and texture features. These production quality gaps reflect resource constraints that prevent exploiters from maintaining professional standards across all communication channels simultaneously. We operationalise findings through a proof-of-concept decision support system providing interpretable risk scores for practitioners. This work demonstrates how rigorous analytical frameworks can address complex humanitarian operations challenges characterised by information asymmetry and limited ground-truth data.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑