Learning to Learn a Language
A 300M-parameter model trained only on synthetic priors learns to predict real languages in context.
Researchers present the Prior-Fitted Language Model (PFLM), a 300-million-parameter byte-level transformer pretrained only on samples from a synthetic non-linguistic prior. With frozen weights and no exposure to real language, it infers a language from a text prefix and predicts its continuation. On Wikipedia in six languages, bits per byte fall from a uniform 8 to between 0.9 and 2.4 at one million bytes of context. Given numerals it approximately counts, compares, and adds, and it compresses source code, speech, and four other non-text domains below gzip and PPMd.
- PFLM is a 300M-parameter byte-level transformer with frozen weights at test time.
- Training uses only synthetic sequences from freshly sampled structural causal models.
- On six-language Wikipedia, bits per byte fall to 0.9-2.4 at one million bytes.
- It also approximates counting and addition and beats gzip and PPMd on six non-text domains.
Full article188 words · extracted from huggingface.co · click to collapse
We present the Prior-Fitted Language Model (PFLM), a 300M-parameter byte-level transformer pretrained only on samples from a synthetic non-linguistic prior. Given a prefix of real text, it learns to predict the language in context with frozen weights, having never seen a word of any real language. Every training sequence is generated by a recurrent structural causal model drawn fresh from a distribution over such models. The model never sees the same language twice during training, so the only way to predict the continuation is to infer the language from the prefix. Samples from this prior share the statistical signatures of natural text: Zipfian frequencies, slow entropy-rate convergence, and long-range dependence. On Wikipedia in six languages, bits per byte fall from the uniform eight to between 0.9 and 2.4 at one million bytes of context. Given numerals instead of text, PFLM learns to count, to compare magnitudes, and to add approximately. It predicts deterministic sequences like Rudin-Shapiro or the prime indicator, and it compresses six non-text domains, from source code to speech, below gzip and PPMd. The model has not learned a language. It has learned to learn one.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.05879