A new research paper explores methods for detecting concept drift in large-scale e-commerce machine learning operations. The study evaluates five multivariate two-sample drift detectors, finding that distributed Maximum Mean Discrepancy with Random Fourier Features on Apache Spark scales effectively. In contrast, the per-dimension Kolmogorov-Smirnov test proved problematic due to statistic saturation with high-cardinality features. The research highlights the challenges of reliable drift detection at scale and the need for future analyses to establish sensitivity bounds. AI
IMPACT Provides scalable methods for maintaining ML model performance in dynamic e-commerce environments.
RANK_REASON Research paper published on arXiv detailing methods for detecting concept drift. [lever_c_demoted from research: ic=1 ai=1.0]
- Apache Spark
- arXiv
- Failing Loudly: An Empirical Study of Methods for Detecting Dataset Shift
- Harvard Dataverse
- Kolmogorov–Smirnov test
- Maximum Mean Discrepancy
- random Fourier features
- Trendyol
AI-generated summary · Google Gemini · from 1 sources. How we write summaries →