Reddit r/LocalLLaMASeptember 22, 2026
Fork of FreeToken with DeepSeek-V4.1, vision and speculative decoding (2x3090 numbers inside)
Excerpt
I've been running FreeToken on my 2x3090 box for a while and ended up maintaining a fork of it. Posting it in case it's useful to anyone else here. Quick context if you haven't used it: FreeToken is an edge-native MoE serving engine. It offloads experts to host RAM/NVMe and co-executes on CPU+GPU so you can run big MoE models on consumer hardware. Upstream is here: https://github.com/FlashML-org/FreeToken What I added on top of upstream: - DeepSeek-V4.1-Flash (mHC, CSA2 sparse attention, lightni