← Back to all articles
arXiv cs.AIAugust 18, 2026

Layers Matter: Why Continual Learning Regularization Should Be Layer-Adaptive

Excerpt

arXiv:2608.15901v1 Announce Type: cross Abstract: Continual learning regularizers like EWC fight forgetting by penalizing changes from previous-task parameters with per-parameter importance, typically diagonal Fisher values. Per-parameter looks more flexible than per-layer, but each layer's diagonal Fisher is a weak summary of its actual curvature, missing the top-eigenvalue information that controls forgetting. Adversarial bit-flip attacks and Hessian-spectrum studies show that this missing per