No AI summary available for this article.
Why It Matters
Reducing the modality gap between image and text representations in CLIP is widely expected to improve cross-modal alignment and downstream performance.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Reducing the modality gap between image and text representations in CLIP is widely expected to improve cross-modal alignment and downstream performance. However, a smaller average image-text gap does not necessarily lead to consistent accuracy gains. We analyze this mismatch from the perspective of the decision structure in zero-shot classification, i.e. selecting the most similar class-text prototype for an input image. Zero-shot accuracy depends not only on average image--text alignment, but also on class-wise decision margins. Using Linear correction as an analytically tractable case, we sh...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.01103v1 · Indexed about 2 hours ago