No AI summary available for this article.
Why It Matters
Open vocabulary 3D semantic segmentation methods typically lift CLIP features into 3D.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Open vocabulary 3D semantic segmentation methods typically lift CLIP features into 3D. This embeds points in a joint vision-language space known to behave like a bag-of-words on compositional tasks. Furthermore, even annotation free variants often require a large 3D training corpus and a dedicated 3D encoder per domain. Instead we use a vision-language model purely as a translator. It produces structured, entity-level descriptions of each posed image. These descriptions are grounded, projected, and aggregated directly in a general-purpose, language-only embedding space, with no 3D training cor...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.09082v1 · Indexed about 1 hour ago