Reddit r/LocalLLaMASeptember 18, 2026
tool-prune: prune 50+ tool schemas down to candidates in 0.4ms (-92% prompt tokens, zero deps)
Excerpt
When running local 7B/8B models with tools, 50+ schemas in context degrades attention and causes distractor hallucinations. Using in-band progressive search (tool_search) costs an entire extra LLM generation turn (+2s), and small models frequently forget to call the search tool or get stuck in discovery loops. tool-prune runs pre-flight in the client before the model is called: const topTools = await router.filter(userPrompt); 0 extra LLM round-trips 0.4ms offline via FWHT (pure JS & Python, 23µ