BrainWideBench: Benchmarking large-scale pretraining and across-animal transfer in multi-region neural recordings
BrainWideBench tests whether pretrained neural representations transfer across mice, behavior, dynamics, and anatomy.
BrainWideBench is a benchmark for across-animal transfer on multi-region neural recordings, built on the International Brain Laboratory Brainwide Map from 139 mice and 276 brain regions. It evaluates whether pretrained representations support behavior decoding, prediction of masked or future neural activity, and recovery of anatomical organization, including finetuning and zero-shot transfer to unseen animals. Pretraining beats matched single-session baselines, but gains depend on alignment between pretraining and downstream tasks, and no method performs uniformly across all three suites.
- Benchmark uses IBL Brainwide Map data from 139 mice across 276 brain regions.
- Three suites test behavior decoding, neural prediction, and anatomical recovery.
- Pretraining beats single-session baselines, but no method wins all suites.
- Zero-shot transfer to unseen animals remains uneven and objective-dependent.
Full article249 words · extracted from arxiv.org · click to collapse
Advances in large-scale neural recording have made it possible to collect data across many animals and distributed brain regions, raising the question of whether this scale can be exploited to learn general-purpose neural representations transferable across diverse downstream tasks. Yet, progress toward this goal has been limited by fragmented evaluation protocols and a narrow focus on individual task domains. Here, we present BrainWideBench, a benchmark for evaluating across-animal transfer on multi-region neural recordings, built on the International Brain Laboratory Brainwide Map dataset of neural and behavioral recordings spanning 276 brain regions from 139 mice performing a sensory-guided decision-making task. The benchmark is organized around three complementary task suites that evaluate whether learned representations support downstream decoding of behavior, can predict masked or future neural activity, and can recover biologically meaningful anatomical organization. With this benchmark, we systematically evaluate pretraining methods across transfer settings, including finetuning on downstream objectives and zero-shot generalization to unseen animals. Our results confirm pretraining improves performance over matched single-session baselines, but we show current methods exhibit heterogeneity in transfer capabilities: gains depend strongly on the alignment between pretraining objectives and downstream tasks. No single approach performs uniformly well across all three suites, and most methods are designed to only address a subset of them. Together, these findings suggest that learning representations that jointly generalize across behavior, dynamics, and anatomy remains an open challenge. By providing a unified and reproducible evaluation suite, BrainWideBench establishes a framework for measuring progress toward general-purpose models of the mouse brain.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.22064