← Back to all articles
Reddit r/LocalLLaMASeptember 9, 2026

Is anyone working on conversation compaction?

Excerpt

In our chat app "harness" we recursively generate summaries, L1 → L2 → L3. L1 summaries are more factual extraction than coherence, then get rolled into a more storytelling L2. Then we keep a tail of always ~10 raw messages with timestamps. We're not coding so this keeps context really clean like 5-8k/tk. However, after like ~120-140 messages, I notice severe degradation in Qwen Flash Next. The model just starts to fall apart, messages quickly become incoherent and comedically strange. But this