No AI summary available for this article.
Why It Matters
A language model normally begins training with random word embeddings: whatever 'banana' means must be learned from training corpora.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
A language model normally begins training with random word embeddings: whatever 'banana' means must be learned from training corpora. I implement St. Augustine's picture of word learning, meaning by ostension, for a small masked language model (DeBERTa) trained on 10M words: before training, visually grounded tokens receive embeddings derived from the image regions they label; other tokens start random. Visual initialization leaves a measurable imprint that lasts until the end of training. At the same time, the effect remains invisible under most BabyLM benchmarks, which probe abstract grammat...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.11870v1 · Indexed about 3 hours ago