No AI summary available for this article.
Why It Matters
Open-vocabulary audio-visual event localization (OV-AVEL) grounds a text-queried event in time from video, audio, and language.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Open-vocabulary audio-visual event localization (OV-AVEL) grounds a text-queried event in time from video, audio, and language. The supervision sources available to this task can differ in temporal-boundary reliability: on OV-AVEBench, our configured visual teacher gives more reliable boundary cues than the configured audio teacher, although the latter is a strong pretrained audio model and remains semantically informative. This is a setting-specific diagnostic rather than a universal ranking of vision and audio. We formulate the resulting challenge as supervision placement: which teacher sign...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.23376v1 · Indexed about 1 hour ago