← Back to all articles
Reddit r/LocalLLaMAAugust 31, 2026

Compact Rollback MTP: a MTP version for QWEN models for those with little vRAM

Excerpt

I've made a modification of llama.cpp MTP for people that want to run models like QWEN 27B on 16GB and similar setup, the focus is reducing the memory cost of MTP allowing more speed for less ctx cost. MTP Mode Maximum Draft (n) Available Context TG (t/s) Standard 2 72,192 39.53 MTP Compact Rollback 5 77,312 46.39 On this example of a (well tuned!) IQ4 running on 16GB you get some +5k ctx and enjoy 17.35% increase on token generation. With MTP the more speculative tokens you generate (n-max) the