arXiv cs.LGOctober 1, 2026
The System Prompt Illusion: How Instruction Preambles Modify Computation in Language Models
Excerpt
arXiv:2609.38205v1 Announce Type: cross Abstract: System prompts are the primary lever practitioners use to control language model behavior, yet what they actually do to the computation inside the transformer remains poorly understood. Across 17 instruction-tuned models spanning 8 architecture families and 1.5B to 72B parameters, we use Centered Kernel Alignment (CKA) to compare layer-wise representations under 20 system prompts in five functional categories. Effects are layer-selective and inst