← Back to all articles
Reddit r/LocalLLaMASeptember 11, 2026

Spomin - Live KV cache compaction (Experimental for Qwen)

Excerpt

I’ve been building "Spomin", a router that replaces context with summaries directly in the KV cache. The goal is to keep long running sessions going without repeatedly stopping for full compaction and reprocessing the context that remains. It uses my llama.cpp fork - https://github.com/alekk89/llama.cpp-kv-surgical-fork for live cache edits. Spomin Router - https://github.com/alekk89/Spomin How it works The router sits between the harness and runtime, preserves the original transcript in chunks,