← Back to all articles
Reddit r/LocalLLaMASeptember 2, 2026

VoxGen, an AMD-optimized TTS inference engine for VoxCPM 2 models

Excerpt

Hi, everyone, I’ve just released VoxGen, a lightweight native inference engine for VoxCPM2, written in Rust and using Vulkan compute instead of Python/PyTorch/CUDA. Why VoxGen? The main reason I started the project was because I needed a decent local text-to-speech solution. I therefore saw VoxCPM 2 as a reasonable solution. However, most frameworks are NVIDIA-first, and VoxCPM 2 is no exception; as a result, my card was severely stuttering, and my GPU was always spiking. Also, having Python and