← Back to all articles
arXiv cs.AIOctober 7, 2026

From Latent Space to Jacobian Space: Measuring, Evading, and Training Against Safety-Content Accessibility

Excerpt

arXiv:2610.04316v1 Announce Type: new Abstract: Output-only safety monitoring sees only the end of a model's computation, yet the model computes its answer before emitting it: what it is internally poised to say is safety-critical. Jacobian-space (J-space) readouts, linear maps from hidden states to the output vocabulary via the model's input-output Jacobian, have been proposed as a window on behaviorally accessible internal content, and a first safety protocol (JADR) showed that danger recognit