arXiv cs.AI / cs.LG / cs.CL·18d agoWhen Does Scale-Invariant Optimization Become Unstable? An Exact Schedule Law with Weight Decay#deep-learning-theory#learning-rate-schedules#normalizationAI research
arXiv cs.AI / cs.LG / cs.CL·18d agoCurriculum Learning as Transport: Understanding Curricula with Wasserstein Geodesics#curriculum-learning#data-ordering#deep-learning-theoryAI research1
arXiv cs.AI / cs.LG / cs.CL·19d agoLatent-MoE: Domain-Aware Mixture-of-Experts for PDEs with Multi-Regime Physics#mixture-of-experts#neural-tangent-kernel#pdesAI research