← Back to all articles
Reddit r/LocalLLaMASeptember 13, 2026

Is there still strong interest in a dense 9b model?

Excerpt

I have a full model, it's ready to train. It's ~9b parameters. 9.4b to be more exact. That includes a 1/2/3 Engram table, Moonshot's AttnRes modeling, and RoPE / NoPE layering at 3:1 as more or less validated by most major labs. It uses the Llama 3 series tokenizer and LM Head as an initial start. The data fed in is logit level extraction from a Llama 3 teaching model. I've already run the first training steps to test that the model is stable, etc. I'm willing to sit and do the pre-IT training o