← Back to all articles
Reddit r/LocalLLaMASeptember 21, 2026

Putting the question before the context took my local Qwen from 89% to 100% on a decision benchmark, and from ~400 ms to ~80 ms

Excerpt

Small, free finding. I use local Qwen models for typed decisions: a state plus a question with fixed allowed answers, and I read the probability of each answer from the logits of one forward pass instead of generating text. I used to build the prompt like you would for a human: situation first, then the question. Flipping it (question and allowed answers first, state last) changed both numbers at once on my benchmark of 38 short situations, Qwen3.6-35B-A3B at 4-bit on an M2 Max: accuracy 89% to