Global Average Precision for Representation Learning
Researchers introduce gSAP, a global average precision loss that can replace InfoNCE in representation learning.
Standard retrieval metrics and losses such as mAP and InfoNCE score each query separately and do not make similarities comparable across queries. Global Average Precision ranks all query-candidate pairs in one list; gSAP is a differentiable surrogate that needs only a similarity matrix and a positive-pair mask. Swapping it into existing recipes improves supervised metric learning, cross-modal alignment, and self-supervised pretraining, retrieving up to four times as many positives at the same precision and improving transfer, kNN, and zero-shot accuracy.
- gSAP is a drop-in ranking loss using the same inputs as existing losses.
- It stays trainable at low temperatures where per-query surrogates lose gradient.
- It retrieves up to four times as many positives at equal precision.
- Models show higher transfer, kNN, and zero-shot classification accuracy.
Full article263 words · extracted from arxiv.org · click to collapse
Standard information retrieval metrics, such as mean Average Precision (mAP), assess performance one query at a time, based on how the similarities between a query and its positives compare against those with its negatives. The same holds for common representation learning losses, such as InfoNCE and per-query AP surrogates. None of them considers whether similarities are comparable across queries, which any system with a single decision threshold relies on. Global Average Precision (gAP) does, by ranking all query-candidate pairs in one list and computing a single AP. We introduce gSAP, a differentiable surrogate of gAP. It needs only a similarity matrix and a binary matrix marking the positive pairs, the same input as existing losses, so it is a drop-in replacement for them and agnostic to the encoder, the modality, and the source of supervision. Since it considers all possible pairwise comparisons in the batch jointly, it also remains trainable at low temperatures, a regime where per-query surrogates run out of gradient. Swapping it into established recipes improves supervised metric learning, cross-modal alignment, and self-supervised pretraining, where, to our knowledge, it is the first ranking loss to replace the community standard InfoNCE in the latter two. Its similarities are more consistent across queries, which drives the gains under a universal threshold. gSAP retrieves up to four times as many positive pairs as the strongest AP surrogate at the same precision, and it degrades the least when queries with no positives in the database are added. Beyond thresholding, models trained with gSAP also learn better representations, with higher transfer, $k$NN and zero-shot classification accuracy.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.09863