← Back to all articles
arXiv cs.LGOctober 7, 2026

Explaining Attention with Program Synthesis

Excerpt

arXiv:2606.19317v3 Announce Type: replace Abstract: A longstanding goal of research on interpretable deep learning is to replace opaque neural computations with human-meaningful symbolic descriptions. In this paper, we propose an approach for approximating the behavior of components of deep networks with executable programs. We focus on attention heads in transformer language models. For a given head, we first compute its associated attention matrices on a collection of randomly selected trainin