← Back to all articles
Reddit r/LocalLLaMAAugust 22, 2026

Sharp template to NInfer: -42% output tokens, same speed

Excerpt

Sharp v22.1 is u/peculiar-ragdoll 's system prompt that makes Qwen answer way more tersely without losing correctness. NInfer is a hyper-tailored inference engine that only runs certain Qwen models on 5090. NInfer doesn't support changing Jinja templates, so I overlaid the behavior in C++ instead in a fork: ninfer-sharp --chat-style sharp-v22.1 appends Sharp's terseness instruction to the system prompt --reasoning-effort with 7 levels, none = thinking off Official model artifact untouched (NInfe