EAServe: Encode-Aware Disaggregated Serving for Multimodal Large Language Models
EAServe raises multimodal LLM serving goodput by controlling encode, prefill placement, and GPU sharing.
EAServe treats the Encode stage as the control point of the Encode-Prefill-Decode pipeline for multimodal LLMs. Its runtime combines load-adaptive micro-batching, rate-controlled partial offload onto a co-resident prefill worker, and dynamic SM partitioning. Hybrid Auto Selection prunes GPU allocations using per-stage capacity profiles and tunes encode batch size and offload ratio with TPE Bayesian optimization. Across image, video, and audio models it reports up to 4.3x goodput versus NVIDIA Dynamo and 1.7x versus vLLM under the same SLOs.
- Per-request encode leaves GPUs idle and starves prefill and decode.
- Runtime adds micro-batching, partial prefill offload, and dynamic SM partitioning.
- HAS profiles capacity, then tunes batch size and offload ratio.
- Up to 4.3x goodput versus NVIDIA Dynamo and 1.7x versus vLLM.
Full article264 words · extracted from arxiv.org · click to collapse
Disaggregating the two stages, Prefill and Decode, onto separate GPU pools is now a standard optimization for (text-only) LLM serving. However, multimodal LLMs (MLLMs), which add a third phase, Encode, pose new challenges for resource allocation. Encode turns images, video, or audio into embeddings that the language model can consume, yielding a three-stage Encode-Prefill-Decode (EPD) pipeline. Existing frameworks offer only partial answers: text-only PD systems lack Encode, while EPD frameworks expose it as a separate service without regulating downstream request flow. The pipeline also carries a structural resource imbalance: every request enters through Encode before downstream work can begin, yet per-request execution leaves the encode GPU severely underutilized even at high loads, starving the downstream Prefill and Decode workers. Addressing this, we reposition Encode as the control point of the EPD pipeline, exposing three tightly coupled dimensions: when work enters downstream, where prefill executes, and how the GPU is shared. We instantiate this in EAServe across two co-designed layers. Its runtime manages load-adaptive micro-batching, rate-controlled partial offload to a co-resident prefill worker, and dynamic SM partitioning for predictable co-location. The configuration layer, Hybrid Auto Selection (HAS), navigates the joint space of GPU allocation, encode batch size, and offload ratio by pruning unbalanced allocations with per-stage capacity profiling and refining the remainder through TPE-based Bayesian optimization. Evaluated on three MLLM architectures spanning image, video, and audio, EAServe delivers up to 4.3x and 1.7x higher goodput than NVIDIA Dynamo and vLLM, respectively, under identical SLO constraints, sustains more balanced and higher GPU utilization across the EPD pipeline, and reaches near-optimal configurations faster than baseline search methods.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.31551