Canonical Color as a Lens into Concept Decodability in Vision Encoders and VLMs
Probing study shows vision encoders make canonical color linearly decodable from grayscale images and tie it to object identity.
Researchers use canonical color as a controlled testbed for measuring conceptual (not just visible) information in vision encoder representations. A dataset of objects with canonical colors was built, and probes on both color and grayscale images show canonical color remains decodable even when color is removed from the input, linked to predicted object identity. Extending to full VLMs, they find post-training has a surprisingly large effect on color decodability in the vision encoder.
- Canonical color remains linearly decodable from grayscale inputs
- Decodable color is tied to predicted object identity, indicating conceptual linkage
- VLM post-training substantially changes color decodability in the vision encoder
- Framework offers a controllable lens for tracing object-level semantics in multimodal models
Full article141 words · extracted from arxiv.org · click to collapse
Visual encoders construct a representation of the image input for Vision-Language models. How much conceptual, as opposed to immediately visible, information does this representation contain? We use canonical color as a controlled test case to ask whether vision encoders make canonical-color information linearly accessible, even when color is removed from the input image. We construct a dataset of objects with canonical colors, and probe vision encoders for both color and object identity using color and grayscale images. We find that canonical color remains decodable from grayscale images, and is tied to predicted object identity, indicating a conceptual link. Extending this analysis to full VLMs, we find that VLM post-training can have a surprisingly large effect on color decodability in the vision encoder. Overall, canonical color provides a usefully controllable lens for tracing object-level conceptual semantic information in vision encoders and VLMs.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.09124