ZeroHour
arXiv cs.CRpublished ()ingested Emily J. Winokur

Nebulon Enterprise Simulated Threats for Phishing Research (NEST-Phish): A Synthetic Enterprise Phishing Email Dataset for Behavioral and Machine-Learning Research

infoResearchimportance 35
AI summary · glm-5.3-flash

Researchers release NEST-Phish, a synthetic enterprise phishing email dataset with matched legitimate and phishing emails and cue annotations for detection research.

Academic researchers introduce NEST-Phish, a publicly released synthetic enterprise phishing email dataset built around a fictitious organization named Nebulon. It contains matched synthetic legitimate and phishing emails across a broad set of workplace communication themes, each with interpretable phishing-cue annotations. Human-subject categorizations and supervised classifier evaluations indicate the dataset supports meaningful variation in phishing judgments and provides learnable signal for detection models. It is intended to support work on phishing detection, human susceptibility, explainability, and benchmark development.

  • Built around fictitious organization Nebulon covering broad workplace communication themes
  • Matched synthetic legitimate/phishing email pairs with interpretable cue annotations
  • Validated through human-subject categorizations and classifier evaluations
  • Targets phishing detection, susceptibility, explainability and benchmark research
VendorsNEST-Phish
OrganizationsNebulon
Full article119 words · extracted from arxiv.org · click to collapse

Phishing remains one of the most persistent cyber threats, yet publicly shareable datasets for studying phishing in realistic enterprise email settings remain limited. To address this gap, we introduce a synthetic enterprise phishing email dataset built around a fictitious organization, Nebulon. The dataset spans a broad set of workplace communication themes and includes matched synthetic legitimate and phishing emails with interpretable phishing-cue annotations. Here, ``legitimate'' denotes the non-phishing class, not legitimately occurring organizational emails. Human-subject categorizations and classifier evaluations show that the dataset supports meaningful variation in phishing judgments while also providing learnable signal for supervised detection. This publicly released resource is intended to support future work on phishing detection, human susceptibility, explainability, and benchmark development in enterprise-like contexts.

Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.04474