No AI summary available for this article.
Why It Matters
In this work, we present a method for shape reconstruction and tracking from video via agentic analysis-by-synthesis.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
In this work, we present a method for shape reconstruction and tracking from video via agentic analysis-by-synthesis. Unlike prior methods which first estimate dense pixel correspondences and then recover object motion from them, our method infers a structured 3D object model, including its geometry and kinematic structure, and uses this model to optimise object track estimates over time. In our optimisation loop, a Vision-Language Model (VLM) agent iteratively refines shape or generalised pose through a render-and-compare loop, combining coarse visual reasoning with numerical pose optimisatio...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.24487v1 · Indexed about 1 hour ago