arXiv cs.CLAugust 19, 2026
SCRIBES: Web-Scale Script-Based Semi-Structured Data Extraction with Reinforcement Learning
Excerpt
arXiv:2510.01832v2 Announce Type: replace Abstract: Semi-structured content in HTML tables, lists, and infoboxes accounts for a substantial share of factual data on the web, yet the formatting complicates usage, and reliably extracting structured information from them remains challenging. Existing methods either lack generalization or are resource-intensive due to per-page LLM inference. In this paper, we introduce SCRIBES (SCRIpt-Based Semi-Structured Content Extraction at Web-Scale), a novel r