← Back to all articles
arXiv cs.LGOctober 7, 2026

Weight Oracles: Reading Neural Network Weights with Language Models

Excerpt

arXiv:2610.07334v1 Announce Type: new Abstract: Interpretability methods for neural networks are predominantly reactive: they analyse activations produced during specific forward passes, requiring known inputs to find hidden capabilities such as backdoors. We propose Weight Oracles, fine-tuned language models that diagnose properties of a target network by reading its raw weights directly, without behavioural testing. We investigate this paradigm in two phases. Phase I establishes feasibility: t