No AI summary available for this article.
Why It Matters
Visual encoders construct a representation of the image input for Vision-Language models.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Visual encoders construct a representation of the image input for Vision-Language models. How much conceptual, as opposed to immediately visible, information does this representation contain? We use canonical color as a controlled test case to ask whether vision encoders make canonical-color information linearly accessible, even when color is removed from the input image. We construct a dataset of objects with canonical colors, and probe vision encoders for both color and object identity using color and grayscale images. We find that canonical color remains decodable from grayscale images, and...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.09124v1 · Indexed about 2 hours ago