← Back to all articles
Reddit r/LocalLLaMASeptember 11, 2026

Thinking that we’ll get safety by CoT traces is wishful thinking. Safety lives in the harness, not the chain of thought

Excerpt

Astra's launch has produced a strange discourse. The reporting that broke the story framed the model's use of recurrent depth primarily as a safety regression, because it means the model reveals less of its "thinking." Spinning latent reasoning as the villain here makes very little sense, especially given all the revelations about problems with CoT transparency and secret message encoding. Because underneath the coverage sits a harmful belief that we can keep the system safe by reading chains of