ALICE: In-context, Zero-shot, Mutual Information Estimation
ALICE estimates mutual information zero-shot from samples without retraining on each distribution.
ALICE is a foundation model trained exclusively on synthetic distributions that estimates mutual information without fitting a new model for each target distribution. Conditioned on samples, it predicts an unseen distribution's rectified-flow velocity field, and mutual information is obtained by integrating the squared difference between joint and conditional fields. On a standard benchmark and on biology, genetics, and neuroscience data it never saw, a single model matches neural estimators trained separately per distribution while handling varied dimensionality and sample sizes.
- Trained only on synthetic distributions, with no per-distribution fitting
- Estimates rectified-flow velocity fields from samples of unseen distributions
- Mutual information comes from a fixed identity on joint versus conditional fields
- Matches specialized neural estimators on a standard benchmark
- Zero-shot results on unseen biology, genetics, and neuroscience data
Full article203 words · extracted from huggingface.co · click to collapse
Estimating mutual information (MI) from samples is a central objective in a variety of scientific fields. Modern neural estimators are accurate in the large-data regime, but they fall short when data is scarce, and each must be fit anew for every distribution under study. Current estimators are moreover tied to specific data types. These constraints limit their adoption in many applications where per-distribution training is impractical and sample sizes are small. We present ALICE, a foundation model that removes per-distribution training, while achieving competitive estimation accuracy. Trained exclusively on a broad family of synthetic distributions, ALICE acts as an in-context estimator of rectified-flow velocity fields: conditioned on samples of an unseen distribution, it estimates that distribution's velocity field without any explicit training. MI is then obtained through a fixed identity that integrates the squared difference between the joint and conditional fields. We validate ALICE on a standard, challenging benchmark and apply it in three domains, biology, genetics, and neuroscience, whose data the model has never seen. For the first time, we show that a single model closes the gap with neural estimators trained separately for each distribution, while natively supporting different data dimensionality and sample cardinality, enabling zero-shot MI analysis across scientific domains.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.34962