Reddit r/LocalLLaMAAugust 22, 2026
Freetokens project is impressive
Excerpt
A new project was released yesterday and I have the opportunity to test it today. Papper: https://arxiv.org/abs/2608.16157 Github: https://github.com/FlashML-org/FreeToken My initial tests with the following setup: RTX 5080 (16 GB) DDR6 64GB AMD Ryzen 9 9950X3D I got 100tok/s on QWEN3.6-35B-A3B NVFP4 (20GB - does not fit in my VRAM). Have you already tried it? (Example bellow with a 1028 token prompt - ~110 tok/s) https://preview.redd.it/x21sl7oo2wkh1.png?width=833&format=png&auto=webp&s=7372da1