Pruned CTC for Memory-Efficient Large-Vocabulary ASR Training
Pruned CTC restricts alignment computation to target tokens and blank, cutting ASR training memory 5.1x for 180K-vocabulary LLMs with minimal accuracy loss.
Pruned CTC restricts CTC alignment computation to the small subset of target tokens plus blank while retaining full-vocabulary normalization, provably equivalent to standard CTC in loss and gradients. With a Zipformer-M encoder and 180K vocabulary, it reduces full-step memory 5.1x with only 17% step-time overhead and matches standard CTC accuracy across three corpora. The authors build LLM-CTC, adapting pretrained LLMs (Qwen3, 0.6B to 32B) for non-autoregressive and streaming ASR, staying within 7% relative WER of LLM-CE while recognizing 7-10x faster.
- Vocabulary pruning is proven exactly equivalent to full-vocabulary CTC in loss and gradients.
- 5.1x memory reduction at 180K vocabulary with only 17% step-time overhead.
- LLM-CTC adapts Qwen3 models 0.6B-32B for streaming ASR within 3-7% relative WER of baselines.
Full article233 words · extracted from huggingface.co · click to collapse
Connectionist temporal classification (CTC) naturally supports offline and streaming speech recognition with utterance-level supervision, but conventional implementations materialize frame-by-vocabulary activations in memory, making CTC training with native LLM vocabularies prohibitively memory-intensive. A key observation is that every valid CTC alignment uses only target tokens and blank, and their union across a batch typically forms a small subset of the full vocabulary. We introduce Pruned CTC, which restricts alignment computation to this subset while retaining full-vocabulary normalization. We prove that this vocabulary reduction is exactly equivalent to full-vocabulary CTC in loss and gradients. Head-and-loss activation memory no longer scales linearly with vocabulary size. We further apply finite-beam alignment pruning. Building on Pruned CTC, we develop LLM-CTC, which adapts pretrained LLMs for non-autoregressive ASR while retaining causal attention and native vocabularies, and extend it to bounded-history streaming, avoiding chunk-level speech--text alignments. Experiments show that, with Zipformer-M encoder and 180K vocabulary, Pruned CTC reduces full-step memory by 5.1times with only 17% step-time overhead. Across three corpora, it matches standard CTC accuracy. On GigaSpeech, across six Qwen3 model sizes from 0.6B to 32B, LLM-CTC remains within 7% relative WER of LLM-CE with 7 to 10times faster recognition; when fine-tuning Qwen3-ASR for bounded-history streaming, LLM-CTC remains within 3% relative WER of matched offline models on the test set. Together, these results establish Pruned CTC as a scalable sequence objective for native-vocabulary LLM ASR across offline and streaming settings.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2609.33645