arXiv cs.CLSeptember 28, 2026
Where a Model Sends Its Own Repeated Token
Excerpt
arXiv:2609.31181v1 Announce Type: new Abstract: Black-box model identification works by scoring a model's response to natural-language prompts. One line of work feeds models a degenerate input -- their own token, repeated -- to find a failure mode rather than an identity. We take that input and ask where the model goes when it does not. For each token t, read argmax p(. | t, t) in one forward pass; the result is a map on the whole vocabulary, with two halves. The first -- which tokens are fixed