arXiv cs.LGOctober 2, 2026
IrekoGPT: Turning Structured Pruning into Post-Hoc Slimmable LLMs
Excerpt
arXiv:2610.00426v1 Announce Type: new Abstract: We introduce IrekoGPT, a post-hoc method for converting pretrained LLMs into slimmable models whose width can be adjusted at inference time. Building on SliceGPT, we retain its projection matrices without pruning them, allowing a single model to expose nested subnetworks at different widths. We improve robustness by calibrating each layer across multiple compression ratios, and correct downstream linear layers through gradient-free ridge regression