← Back to all articles
arXiv cs.CLSeptember 11, 2026

LILA: Calibration-Free Structured Pruning of Large Language Models via Latent Spectral Geometry

Excerpt

arXiv:2609.11163v1 Announce Type: cross Abstract: Structured pruning of large language models (LLMs) offers hardware-efficient compression, yet existing methods require calibration data, gradient computation, or large auxiliary policy networks at pruning time. LILA (\emph{Latent-Informed Layer Analysis}) scores neuron importance via the Kolmogorov--Smirnov (KS) distance between empirical singular value distributions of the full and neuron-ablated feed-forward network (FFN) weight matrix, providi