← Back to all articles
arXiv cs.CLSeptember 24, 2026

Six Layers Less: Encoder Pruning for Whisper with Label-Free Recovery

Excerpt

arXiv:2609.27980v1 Announce Type: new Abstract: Pruning large pre-trained transformer-based ASR models such as OpenAI's Whisper has seen great adoption, as pruning the decoder led to significant end-to-end transcription speedups. For instance, the {\tt whisper-large-v3-turbo} variant reduced the decoder from 32 to 4 layers, while Distill-Whisper similarly reduced the decoder to only 2 layers. Although some attention has been put towards reducing the size of the encoder, no approach has seen wide