← Back to all articles
Reddit r/LocalLLaMASeptember 22, 2026

Model grafting: turning Qwen3.5-4B into a causal encoder-decoder after the fact

Excerpt

Recently, the new DeepSeek-V4.1-Flash architecture showed how a causal encoder-decoder can work, but it was trained from scratch. Model Grafting does it to an existing model: cut at some depth, let the lower layers read the prompt, and use the upper layers get for encoder's residual stream as prefix KV via identity-init adapters, then heal with self-distillation from the unmodified parent. Decoding part stays the same, this method was described in this blog post https://latentnode.pages.dev/arti