← Back to all articles
arXiv cs.CLSeptember 23, 2026

Diffusion Drafts, AR Verifies: Accelerating Document OCR with Self-Speculative Decoding

Excerpt

arXiv:2609.26638v1 Announce Type: new Abstract: Autoregressive OCR vision-language models accurately convert document images into text and structured markup, but require one sequential decoding step per output token, limiting inference speed. Unlike open-ended text generation, OCR outputs are strongly grounded in the input image, making diffusion-based parallel generation promising. However, when several tokens are predicted in one diffusion step, each is predicted before the others are known. C