← Back to all articles
Reddit r/LocalLLaMAAugust 31, 2026

New beellama fork 76% faster tg with kvarn KV quants

Excerpt

When using kvarn quants at low context depth tg speed is similar to llama.cpp on and equivalent qx_x quant. However, as context depth grows kvarn tanks your tg speed. This fork optimises kvarn to have similar or better performance ay high context depths than llama.cpp at an equivalent qx_x quant and in my testing up to 76% faster tg than beellama's implementation of kvarn. My testing capped at ctx 99328 but for higher context your gains will be even better. From the github (translated from Russi