SynCo: Learning Cross-Modal Synergy by Contrasting Interaction Residuals
SynCo trains multimodal synergy with contrastive loss on interaction residuals, gaining 5.98 points on Trifeature.
SynCo targets undertrained synergy in multimodal contrastive learning, using Partial Information Decomposition’s split of redundancy, uniqueness, and synergy. A linear projector predicts the fused representation from unimodal features, and the interaction residual—what is not linearly predictable from either modality—receives its own contrastive loss at low extra cost. On Trifeature, SynCo reports a 5.98 percentage-point synergy gain over the baseline, and it matches or outperforms prior methods on MultiBench, DARai, and MM-IMDb without changing the fusion architecture.
- SynCo applies contrastive supervision to an interaction residual to learn multimodal synergy.
- A linear projector removes the component predictable from unimodal features alone.
- Trifeature synergy rises 5.98 points; MultiBench, DARai, and MM-IMDb match or beat priors.
Full article195 words · extracted from arxiv.org · click to collapse
Multimodal contrastive learning is a dominant paradigm for learning transferable representations from unlabeled data, but standard objectives primarily capture information that is redundant between modalities. Partial Information Decomposition (PID) shows that task-relevant information in multimodal data decomposes into three components: redundancy shared between modalities, uniqueness specific to each modality, and synergy available only from their joint observation. Recent frameworks extend contrastive learning to capture all three components, yet synergy remains undertrained in practice. We propose SynCo (Synergy Contrastive Learning), a method that directly addresses synergy undertraining through dedicated supervision on an interaction residual. SynCo fits a linear projector to predict the fused representation from independently computed unimodal features, and the resulting interaction residual, which removes the linearly unimodal-predictable component, receives dedicated contrastive supervision at negligible computational cost. On the controlled Trifeature benchmark, SynCo achieves state-of-the-art synergy capture with a $+5.98\%$ gain over the baseline, and on real-world benchmarks from MultiBench, DARai, and MM-IMDb, SynCo consistently outperforms or matches prior methods across diverse modality combinations and task types. The method operates as a plug-in to existing contrastive multimodal frameworks without modifying the underlying fusion architecture and can further improve synergy capture when combined with other methods.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.32846