← Back to all articles
arXiv cs.CLSeptember 11, 2026

Perturbation: A simple and efficient adversarial tracer for representation learning in language models

Excerpt

arXiv:2603.23821v2 Announce Type: replace Abstract: Linguistic representation learning in deep neural language models (LMs) has been studied for decades, but finding representations in LMs remains an unsolved problem. On the one hand, unconstrained alignments may trivialize the notion of representation (Sutter et al., 2025); on the other, even recently popularized linear approaches may not always be faithful to natural model behavior (Arora et al. 2024). Here we escape this dilemma by reconceptu