Telescopic Language Models
Telescopic Language Models train one nested Transformer that is valid at every layer-prefix depth.
Telescopic Language Models are nested-capacity Transformers trained with stochastic prefix supervision and a full-capacity anchor so every depth prefix is a valid language model. Each step uses two forward-backward passes, with no architectural change and nothing extra at inference. On a 200M proxy trained on 20B FineWeb-Edu tokens, one run was valid at all twenty layer prefixes and reduced area under the quality-budget curve by 43-44% versus fixed-exit Matryoshka suites, matching them at full capacity at about 12% lower GPU cost. Fixed-exit baselines remained near chance at intermediate depths that were not directly supervised.
- Stochastic prefix supervision makes every layer prefix a valid language model.
- Two passes per step; no architecture change or extra inference cost.
- A 200M model on 20B FineWeb-Edu tokens is valid at all 20 prefixes.
- Quality-budget area falls 43-44% versus fixed-exit Matryoshka suites.
Full article253 words · extracted from arxiv.org · click to collapse
One deployed language model must often serve many compute budgets, yet serving each budget still means a separate training or compression run per point. We train a Telescopic Language Model (TLM) to be that continuum: a nested-capacity Transformer supervised by stochastic prefix supervision with a full anchor. At every step, one randomly truncated prefix of the capacity axis is trained against the full next-token target, alongside one full-capacity pass, so the trained artifact is a valid language model at every depth. Two forward-backward passes per step, no architectural change, nothing extra at inference. Fixed-exit suites such as Matryoshka Language Model Suites (MLMS) occupy one point in this design space, and the point has a cost: supervising only a few fixed exits leaves the nested model at chance level everywhere else (perplexity 10^2-10^5 in our baselines). On a 200M proxy suite (20B FineWeb-Edu tokens, identical data stream for all methods), a single TLM run is a valid language model at every one of its twenty layer prefixes, in perplexity and on perplexity-sensitive downstream tasks, reducing the area under the quality-budget curve by 43-44% relative to the fixed-exit suites while matching them at full capacity, at ~12% lower GPU cost per run. The prefix sampling density is a dial: concentrating it on a few depths recovers fixed-exit quality there at the price of the continuum, so the operating points become a training-time choice rather than an architectural one. These results indicate that the training objective, not the nesting itself, is what makes a model elastic.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.35769