Beyond Marginal Monitoring: Distributed Joint-Distribution Testing for Data Concept Drift in Large Scale E-Commerce Operations
Large-scale tests find Spark MMD reliably detects strong concept drift in a 137.5-million-row table.
The paper evaluates five multi-column two-sample tests for concept-drift detection on large e-commerce feature tables. Environments include the Harvard Dataverse, a Failing Loudly reproduction with mean absolute error between 0.030 and 0.053, and synthetic drift injected into Trendyol’s 137.5-million-row collection-ranking table. Distributed Maximum Mean Discrepancy with Random Fourier Features on Apache Spark achieves a Pearson correlation of 0.940 with expected drift magnitude in the strong regime, with an 80.4% true positive rate and 3.2% false positive rate. The per-dimension Kolmogorov-Smirnov test saturates on ID-like columns, and detectors struggle when realized-flip fractions are at most 0.57%.