Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching
Continuous depth batching enables efficient depth-adaptive inference for looped language models, reaching up to 99% of estimated maximum speedup.
The paper introduces continuous depth batching, a method for efficient depth-adaptive inference in looped language models that cannot use uniform batching systems such as vLLM. It rebuilds batches between loop steps, schedules looped and non-looped layers, manages looped KV caches, and predicts exits so batches can be prepared asynchronously. Experiments on Ouro 1.4B and Huginn 3.5B find fully looped architectures best suited to the method, because large non-looped layers slow scheduling. Continuous depth batching reaches up to 99% of the estimated maximum speedup.
- Forms new batches between loop steps for tokens with different depths
- Predicts which tokens exit loops so batches can be prepared asynchronously
- Experiments use Ouro 1.4B and Huginn 3.5B looped language models
- Achieves up to 99% of the estimated maximum available speedup
Full article191 words · extracted from huggingface.co · click to collapse
A main promise of looped language models is depth-adaptive inference. By looping a block of shared layers a variable number of times, the model can use less compute for "easy" tokens and more for "hard" ones. However, tokens with different numbers of loops cannot share a uniform forward pass and therefore cannot be handled by standard batching systems such as vLLM. The practical value of depth-adaptive inference thus hinges on whether batching can be made efficient. We introduce the first efficient method for depth-adaptive looped LMs via continuous depth batching (CDB), which forms new batches between loop steps. Our method dynamically schedules looped and non-looped parts of the architecture, manages looped KV-caching, and predicts which tokens will exit the loop in advance so it can prepare batches asynchronously. Experiments on Ouro 1.4B and Huginn 3.5B show that fully looped architectures are best suited to depth-adaptive inference, as large non-looped layers outside the recurrent core (e.g., token embedding, LM head, and unshared transformer blocks) slow down and complicate scheduling. Overall, CDB realizes up to 99% of the estimated maximum speedup available, leaving further gains primarily dependent on model architecture and exit behavior.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2608.09444