WaveFront Decoding: Parallelized Self-Speculative Decoding for Looped Language Models
Wavefront Decoding speeds looped models Ouro-2.6B and Huginn-3.5B by up to 4.81x over autoregressive decoding.
Wavefront Decoding is a training-free self-speculative decoding method for looped language models that repeatedly apply a weight-shared block. It uses intermediate recurrence outputs as drafts and batches token states at different positions and depths in one recurrent-block call, drafting new positions while verifying earlier ones. Across six Spec-Bench categories it reports 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B versus autoregressive decoding, rising to 4.81x on Huginn-3.5B with cross-recurrence KV sharing. Code is released.
- Wavefront Decoding is training-free self-speculative decoding for weight-shared looped language models.
- Mixed-depth token states are batched in one recurrent-block call along a diagonal wavefront.
- Drafting and verification run together; rejected drafts are corrected with full-depth predictions.
- Speedups are 2.42x on Ouro-2.6B and up to 4.81x on Huginn-3.5B with cross-recurrence KV sharing.
Full article174 words · extracted from huggingface.co · click to collapse
Looped language models repeatedly apply a weight-shared block to increase effective depth without increasing parameter count, but the resulting T sequential recurrent-block calls per generated token substantially increase decoding latency. To address the issue, we introduce Wavefront Decoding (WFD), a training-free self-speculative decoding framework designed for looped language models. WFD exploits two properties of these architectures: intermediate recurrence outputs provide effective draft predictions, and weight sharing allows token states at different positions and recurrence depths to be processed in one batched recurrent-block call. WFD organizes these mixed-depth states into a diagonal wavefront, continuously drafting new positions at shallow depth while advancing earlier positions toward full-depth verification. Unlike the phase-separated draft-then-verify schedule, WFD therefore concurrently batches drafting and verification within the same recurrent calls, while rejected drafts are corrected using full-depth predictions. Across six Spec-Bench task categories, WFD achieves 2.42x speedup on Ouro-2.6B and 3.54x on Huginn-3.5B over autoregressive decoding, consistently outperforming draft-then-verify. Cross-recurrence KV sharing further reduces wavefront KV traffic and increases WFD's speedup to 4.81x on Huginn-3.5B. The code is available at https://github.com/summerbro-hhj/wavefront-decoding.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.23033