Beyond Marginal Monitoring: Distributed Joint-Distribution Testing for Data Concept Drift in Large Scale E-Commerce Operations
Large-scale tests find Spark MMD reliably detects strong concept drift in a 137.5-million-row table.
The paper evaluates five multi-column two-sample tests for concept-drift detection on large e-commerce feature tables. Environments include the Harvard Dataverse, a Failing Loudly reproduction with mean absolute error between 0.030 and 0.053, and synthetic drift injected into Trendyol’s 137.5-million-row collection-ranking table. Distributed Maximum Mean Discrepancy with Random Fourier Features on Apache Spark achieves a Pearson correlation of 0.940 with expected drift magnitude in the strong regime, with an 80.4% true positive rate and 3.2% false positive rate. The per-dimension Kolmogorov-Smirnov test saturates on ID-like columns, and detectors struggle when realized-flip fractions are at most 0.57%.
- Five multivariate two-sample drift tests compared at industrial scale.
- Benchmark uses Trendyol’s 137.5-million-row ranking feature table.
- Spark MMD with random Fourier features reaches r = 0.940 in the strong regime.
- Reported true positive rate is 80.4% and false positive rate 3.2%.
- Per-dimension Kolmogorov-Smirnov saturates on ID-like columns.
Full article197 words · extracted from arxiv.org · click to collapse
Concept drift threatens production machine learning, yet the empirical behavior of multivariate two-sample drift detectors at scale remains under-characterized. Existing benchmarks rarely address the hundreds of millions of rows and high-cardinality features typical of industrial-operational datasets. We evaluate five multi-column two-sample tests (marginal, projection-based, and kernel embedding methods) across three complementary environments: the Harvard Dataverse, a validated Failing Loudly reproduction (mean absolute error between 0.030 and 0.053), and a novel synthetic-injection benchmark on the 137.5-million-row Trendyol collection-ranking feature table. Testing four drift types across two severity-scope regimes, we demonstrate that distributed Maximum Mean Discrepancy with Random Fourier Features on Apache Spark scales robustly. Averaged over the four drift types in the strong regime and under a calibrated threshold, it achieves a Pearson correlation of r = 0.940 with expected drift magnitude, an 80.4% true positive rate, and a 3.2% false positive rate. Conversely, the per-dimension Kolmogorov-Smirnov test failed due to statistic saturation from ID-like columns under asymmetric sampling, establishing a critical constraint for large-scale sampling design. At weak configurations (realized-flip fractions of at most 0.57%), detectors struggled to reliably discriminate, highlighting the need for future intensity-grid power analyses to distinguish fundamental sensitivity bounds from scalable threshold shifts.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.08132