arXiv cs.CLSeptember 23, 2026
ARAFA: An LLM-Generated Arabic Fact-Checking Dataset
Excerpt
arXiv:2609.25833v1 Announce Type: new Abstract: Automatic fact-checking poses a significant challenge in Arabic natural language processing due to the scarcity of datasets and resources. In this manuscript, we introduce Arafa, a new large-scale dataset for fact-checking in Modern Standard Arabic, constructed through an automated framework leveraging large language models (LLMs). The dataset was constructed through a three-step pipeline: (1) claim generation from Arabic Wikipedia pages with suppo