← Back to all articles
arXiv cs.CLSeptember 22, 2026

MechaTerp-TRACE: A Novel Approach for Component Ablation Analysis in Language Models

Excerpt

arXiv:2609.22163v1 Announce Type: new Abstract: Interpretability research on large language models has produced accounts of factual recall in feed-forward layers and of token relationships in self-attention, but little work offers a unified way to compare the causal contribution of different architecture components to a model's output. We introduce MechaTerp (the Mechanistic Interpretability suite) -TRACE (subset for Teacher-forced Registry of Ablated Component Effects), an architecture and stud