Reddit r/LocalLLaMASeptember 21, 2026
You can use any LLM just like JEV
Excerpt
You can simply run any GGUF with llama.cpp with n_predict=1 and n_probs=10, disable reasoning, and prompt it such as "If the following email is spam, respond with 1, if not spam, respond with 0. Do not respond with anything other than 1 or 0. Email: ...." And that is it! It returns confidence percentages such as: 1 = 94.9% 0 = 5.08% Example: llama-server -m "C:\Users\MyUserName\llama.cpp\models\Spark-X2.5-4B-Q4_K_M.gguf" -c 4096 -ngl all -fit off -fa on -b 2048 -ub 512 -np 1 --cache-ram 0 --reas