← Back to all articles
arXiv cs.CLSeptember 21, 2026

Evaluating In-Context Learning and Retrieval Strategies for Devanagari Post-OCR Correction

Excerpt

arXiv:2609.21595v1 Announce Type: new Abstract: In-context learning using Large Language Models (LLMs) offers a compelling path to training-free post-OCR correction, yet its effectiveness for Devanagari script remains entirely unexplored. We present the first systematic evaluation of LLMs (3B-32B) for post-OCR correction in Hindi and Marathi, comparing three in-context example retrieval strategies: domain-random selection, dense semantic retrieval, and our proposed CharBM25, which retrieves exam