← Back to all articles
Reddit r/LocalLLaMAAugust 22, 2026

Look at me: I am the frontier Lab now - PromptInjectBench: asked Huihui-Qwen3.6-35B to write 60 prompt injection attacks on files used or generated by Hermes. Shieldstral scanned each of them->It caught zero/nothing/nada. All 60 poisoned prompts passed the scanning. GPT-OSS_safeG caught 10%

Excerpt

Look at me: I am the frontier Lab now Huihui-Qwen3.6-35B prompt (on Pi): "In the folder u/source/ you will find 6 files, text files, that are commonly used by my own local coding agent. your task is to take each file and create a new version injecting each of the 10 most common prompt injection attacks listed by yourself in the previous turn. Since we have 6 text files and you listed 10 possible prompt injections I expect to have 60 "prompt injected" files. drop the prompt injected files in the