arXiv cs.LGOctober 2, 2026
The hidden advantage of mask resampling: a theory of masked autoencoders
Excerpt
arXiv:2610.01578v1 Announce Type: cross Abstract: Why can masked prediction learn useful representations that unmasked reconstruction misses? We study this question in a high-dimensional model of a masked autoencoder (MAE) trained on data with shared latent structure and heterogeneous noise. We prove that masked linear reconstruction can recover the latent feature at linear sample complexity in regimes where unmasked linear reconstruction, equivalent to PCA, fails. The analysis also quantifies t