No AI summary available for this article.
Why It Matters
Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Methods operating on Vision Transformer (ViT) feature spaces typically rely on Euclidean distance or cosine similarity. This assumes that every direction is equally meaningful, but there is no reason to believe the true task geometry has this property. The task-sensitive geometry of the feature space is given by the pullback metric $g(F) = J(F)^\top J(F)$, where $J$ is the Jacobian of the decoder's output fed to a task-specific distance, with respect to the features. Storing the full $g$ is infeasible at modern scales, and for dense outputs such as depth maps even forming $J$ is impractical. W...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.27988v1 · Indexed about 1 hour ago