← Back to all articles
Reddit r/MachineLearningAugust 23, 2026

28 TPS on Qwen2.5-7B across two separate cloud regions over public WAN using speculative decoding + CUDA Graphs [P]

Excerpt

been building ShardFlow for the past few months, a distributed LLM inference framework that splits any HuggingFace transformer across N GPU machines and uses neural speculative decoding to deal with WAN latency. the setup for the benchmark: two T4 nodes in separate GCP regions (Iowa + Oregon) talking through an AWS EC2 TCP relay in Ohio. ~86ms RTT on public internet. the key insight with speculative decoding here is that WAN latency stops being a per-token cost and becomes a per-round cost. with