Joint and Cross-Modal Video-Audio Generation and Editing: A Unified Formulation and Design Taxonomy
A survey unifies joint, cross-modal, and coupled video-audio generation and editing under one taxonomy.
The paper reviews methods that generate or edit video and audio jointly or across modalities, focusing on temporal and semantic coherence. It casts joint generation, cross-modal generation, and joint editing as three problems on one distribution over audio-visual pairs. A taxonomy compares methods on five design axes and maps joint audio-visual editing to nine categories spanning 28 edit types. It also covers datasets, metrics, and open problems.
- Joint, cross-modal, and editing tasks share one audio-visual distribution.
- Taxonomy compares methods along five design axes.
- Joint editing maps to nine categories and 28 edit types.
- Survey covers methods, datasets, metrics, and open problems.
Full article128 words · extracted from huggingface.co · click to collapse
Video and audio are perceived together, yet most generative models treat them in isolation. We examine methods that model the two modalities jointly, generate one from the other, or edit them in a coupled manner, organized around a single question: how is the output kept coherent across modalities in time and semantics? A unified formulation casts joint generation, cross-modal generation, and joint editing as three problems defined on a single distribution over audio-visual pairs, and a taxonomy compares methods along five design axes. To our knowledge, this is the first overview to systematically taxonomize joint audio-visual editing, which we map as nine edit categories spanning 28 edit types. We describe methods, datasets, and metrics for each setting and close with the open problems we view as most consequential.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.34381