Detecting Deceptive Recruitment: A Signal-theoretic Machine Learning Framework for Early Identification of Labour Exploitation
Researchers build a multimodal ML framework detecting deceptive job advertisements used for labour exploitation, achieving ROC-AUC 0.87–0.97 across individual modalities.
Using 464 verified cases (164 deceptive, 300 legitimate) gathered through anti-slavery charities across nine origin countries and 21 industries, researchers combined computer vision, NLP, and semantic embeddings to flag exploitative recruitment ads. SHAP analysis identified text quality, readability indices, risk keyword density, and visa sponsorship mentions as top discriminators, with combined modalities reaching ROC-AUC up to 0.97. Findings are operationalized in a proof-of-concept decision support system providing interpretable risk scores for practitioners.
- 464 verified deceptive and legitimate job ads from nine origin countries and 21 industries
- Multimodal CV/NLP/embedding models achieve ROC-AUC 0.87–0.97
- SHAP shows text quality and risk-language keywords are primary discriminators
- Proof-of-concept tool provides interpretable risk scores for anti-slavery practitioners
Full article203 words · extracted from arxiv.org · click to collapse
Deceptive online job advertisements have emerged as a primary pathway into forced labour, yet systematic detection methods remain underdeveloped due to data scarcity and absence of empirically validated indicators. We formalise this detection challenge as a classification problem under signalling theory, where exploiters transmit costless signals mimicking legitimate communications across textual, visual, and structural dimensions. Using 464 verified cases (164 deceptive, 300 legitimate) collected through anti-slavery charities across nine origin countries and 21 industries, we develop multimodal detection models combining computer vision, natural language processing, and semantic embeddings. Through systematic feature ablation experiments and repeated stratified cross-validation, we demonstrate that individual modalities achieve substantial discriminatory power (ROC-AUC: 0.87--0.97), whilst their integration yields modest further gains. SHAP-based analysis reveals that text quality and domain-specific risk language are the primary discriminators, with readability indices, risk keyword density, and visa sponsorship mentions ranking highest, followed by visual colour and texture features. These production quality gaps reflect resource constraints that prevent exploiters from maintaining professional standards across all communication channels simultaneously. We operationalise findings through a proof-of-concept decision support system providing interpretable risk scores for practitioners. This work demonstrates how rigorous analytical frameworks can address complex humanitarian operations challenges characterised by information asymmetry and limited ground-truth data.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.20336