← Back to all articles
arXiv cs.CLSeptember 18, 2026

LMEnt: A Suite for Analyzing Knowledge in Language Models from Pretraining Data to Representations

Excerpt

arXiv:2509.03405v2 Announce Type: replace Abstract: Language models (LMs) increasingly drive real-world applications that require world knowledge. However, the internal processes through which models turn data into representations of knowledge and beliefs about the world are poorly understood. To facilitate such studies, we present LMEnt, a suite including (1) a knowledge-rich pretraining corpus, fully annotated with entity mentions based on Wikipedia, (2) an entity-based retrieval method over p