New research targets efficient LLM inference: depth-adaptive batching for looped models and encode-aware serving for multimodal LLMs
Two papers published a day apart — continuous depth batching for looped language models and EAServe for multimodal LLM serving — report large efficiency gains: up to 99% of estimated maximum speedup and up to 4.3x goodput over NVIDIA Dynamo, respectively.
Two independent research papers on efficient LLM inference appeared on consecutive days. The first, surfaced via Hugging Face daily papers on 2026-09-24, presents continuous depth batching, a method for depth-adaptive inference in looped language models that cannot use uniform batching systems such as vLLM. It rebuilds batches between loop steps, schedules looped and non-looped layers, manages looped KV caches, and predicts which tokens exit loops so batches can be prepared asynchronously. Experiments on Ouro 1.4B and Huginn 3.5B found fully looped architectures best suited to the method, because large non-looped layers slow scheduling, and the approach reaches up to 99% of the estimated maximum speedup. The second paper, posted to arXiv (cs.AI, cs.LG, cs.CL) on 2026-09-25, introduces EAServe, which treats the Encode stage as the control point of the Encode-Prefill-Decode pipeline for multimodal LLMs, noting that per-request encode leaves GPUs idle and starves prefill and decode. Its runtime combines load-adaptive micro-batching, rate-controlled partial offload onto a co-resident prefill worker, and dynamic SM partitioning, while Hybrid Auto Selection prunes GPU allocations using per-stage capacity profiles and tunes encode batch size and offload ratio with TPE Bayesian optimization. Across image, video, and audio models, EAServe reports up to 4.3x goodput versus NVIDIA Dynamo and 1.7x versus vLLM under the same SLOs. The two reports describe separate papers and do not conflict.
- Continuous depth batching paper surfaced on Hugging Face daily papers on 2026-09-24; EAServe appeared on arXiv (cs.AI, cs.LG, cs.CL) on 2026-09-25.
- Continuous depth batching enables depth-adaptive inference for looped language models that cannot use uniform batching systems such as vLLM.
- The method rebuilds batches between loop steps, schedules looped and non-looped layers, manages looped KV caches, and predicts which tokens exit loops so batches can be prepared asynchronously.
- Experiments used the Ouro 1.4B and Huginn 3.5B looped language models; fully looped architectures were found best suited because large non-looped layers slow scheduling.
- Continuous depth batching reaches up to 99% of the estimated maximum available speedup.
- EAServe treats the Encode stage as the control point of the Encode-Prefill-Decode pipeline for multimodal LLMs.
- EAServe's runtime combines load-adaptive micro-batching, rate-controlled partial offload onto a co-resident prefill worker, and dynamic SM partitioning.
- Hybrid Auto Selection prunes GPU allocations using per-stage capacity profiles and tunes encode batch size and offload ratio with TPE Bayesian optimization.
Coverage timelineoldest first · each row is one article
- · 5d agoDepth-adaptive Inference of Looped Language Models via Continuous Depth Batching
Hugging Face daily papers· 48
Continuous depth batching enables efficient depth-adaptive inference for looped language models, reaching up to 99% of estimated maximum speedup.
- · 4d agoEAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models
arXiv cs.AI / cs.LG / cs.CL· 48
EAServe raises multimodal LLM serving goodput by controlling encode, prefill placement, and GPU sharing.