arXiv cs.CLSeptember 11, 2026
Augustinian BabyLM: What Ostensive Definition Can and Cannot Teach a Small Language Model
Excerpt
arXiv:2609.11870v1 Announce Type: new Abstract: A language model normally begins training with random word embeddings: whatever 'banana' means must be learned from training corpora. I implement St. Augustine's picture of word learning, meaning by ostension, for a small masked language model (DeBERTa) trained on 10M words: before training, visually grounded tokens receive embeddings derived from the image regions they label; other tokens start random. Visual initialization leaves a measurable imp