ZeroHour
arXiv cs.AI / cs.LG / cs.CLpublished ()ingested Jinting Wang

CMA-OT: Hierarchical Expert Supervision for Dance-to-Music Generation

infoAI researchimportance 5
AI summary · glm-5.3

CMA-OT aligns a music generator's latent features with hierarchical expert representations via curriculum learning and scale-aware optimal transport, improving dance-to-music quality.

CMA-OT introduces curriculum-guided multi-scale representation alignment with scale-aware optimal transport for dance-to-music generation. An external music expert provides hierarchical supervision over the generator's latent features, progressively transferring musical knowledge for stable representation learning. The optimal transport mechanism handles temporal mismatch and semantic variation across expert scales. Experiments on two datasets show state-of-the-art rhythmic synchronization, perceptual quality, and overall music generation.

  • External music expert supervises generator latent features hierarchically
  • Curriculum strategy progressively transfers musical knowledge during training
  • Scale-aware optimal transport aligns features under temporal mismatch
Full article208 words · extracted from arxiv.org · click to collapse

Dance-to-music (D2M) generation aims to synthesize music that is rhythmically and stylistically aligned with dance videos. A key challenge arises from the semantic mismatch between sparse dance cues, such as rhythm and style, and the dense information required for music composition, including structure, instrumentation, and expressive dynamics. Existing methods typically rely on these sparse cues and supervise only the final audio output, resulting in poorly learned music representations and generated music with limited musicality and structural coherence. To address these issues, we propose Curriculum-guided Multi-scale representation Alignment with scale-aware Optimal Transport (CMA-OT), a novel paradigm that leverages an external music expert to provide hierarchical supervision for the generator's latent features, bridging the semantic gap and enhancing representation learning. To effectively incorporate hierarchical supervision, we introduce a curriculum-guided multi-scale learning strategy that progressively transfers musical knowledge from the expert to the music generator, enabling stable and effective representation learning. Moreover, to accommodate the semantic and structural variations across different expert scales and achieve fine-grained alignment under temporal mismatch, we propose a scale-aware optimal transport alignment mechanism, which models soft correspondences between hierarchical expert representations and the generator's latent features. Extensive experiments on two datasets demonstrate that CMA-OT achieves state-of-the-art performance in rhythmic synchronization, perceptual quality, and overall music generation.

Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.13118