← Back to all articles
Reddit r/LocalLLaMASeptember 9, 2026

DeepSeek-V4-Flash-Vision-Exp (285B MoE) on 10-12x RTX 3090 — spec decoding, vision

Excerpt

Running the full deepseek-ai/DeepSeek-V4-Flash-Vision-Exp on consumer Ampere — 10-12x RTX 3090, SM86-compatible vLLM build. 285B MoE, FP4 experts + FP8 attention, 157 GB weights. Highlights: - **60+ tok/s** decode, DSpark spec (k=3) on 10 GPUs (TP2xPP5), at a 240 W cap - **120+ tok/s** on 12 GPUs (TP4xPP3) - **Vision + spec + tool calls all working** - **1M context** (no offload) / **4M** (RAM offload) - ~3,500 tok/s long-context prefill Fully documented + reproducible: - Pre-built image: `docke