Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Pruned CTC restricts alignment computation to target tokens and blank, cutting ASR training memory 5.1x for 180K-vocabulary LLMs with minimal accuracy loss.
Pruned CTC restricts CTC alignment computation to the small subset of target tokens plus blank while retaining full-vocabulary normalization, provably equivalent to standard CTC in loss and gradients. With a Zipformer-M encoder and 180K vocabulary, it reduces full-step memory 5.1x with only 17% step-time overhead and matches standard CTC accuracy across three corpora. The authors build LLM-CTC, adapting pretrained LLMs (Qwen3, 0.6B to 32B) for non-autoregressive and streaming ASR, staying within 7% relative WER of LLM-CE while recognizing 7-10x faster.