ZeroHour
arXiv cs.AI / cs.LG / cs.CLpublished ()ingested Hongbo Chen1

General Quantification of Covariate and Concept Shifts

infoAI researchimportance 28
AI summary · glm-5.3-flash

Paper proposes γ*-concept shifts via entropic optimal transport, deriving estimable generalization bounds unifying covariate and concept shift under distribution shift.

The authors show existing definitions of concept shift break when source and target supports mismatch and propose γ*-concept shifts grounded in entropic optimal transport. They derive a general error bound covering broad loss functions, label spaces and stochastic labeling, plus estimators with concentration guarantees. The resulting DataShifts algorithm quantifies distribution shifts and estimates the error bound in most applications, addressing learning bounds that were previously non-estimable from samples.

  • Introduces γ*-concept shifts defined through entropic optimal transport.
  • Unifies covariate and concept shift in a single general error bound.
  • Provides estimators with concentration guarantees and the DataShifts quantification algorithm.
ProductsDataShifts
Full article130 words · extracted from arxiv.org · click to collapse

Generalization under distribution shift remains a core challenge in modern machine learning, yet existing learning bound theory is limited to narrow, idealized settings and is non-estimable from samples. In this paper, we bridge the gap between theory and practical applications. We first show that existing definition of concept shift breaks when the source and target supports mismatch. Leveraging entropic optimal transport, we propose a key notion: $γ^{*}\!$-concept shifts, and derive a general error bound unifying covariate and $γ^{*}\!$-concept shifts, which applies to broad loss functions, label spaces, and stochastic labeling. We further develop estimators for these shifts with concentration guarantees, and the DataShifts algorithm, which can quantify distribution shifts and estimate the error bound in most applications - a rigorous and general tool for analyzing learning error under distribution shift.

Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.11918