No AI summary available for this article.
Why It Matters
Instruction-guided music editors typically process each request independently, limiting their ability to support workflows in which users progressively refine a track.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Instruction-guided music editors typically process each request independently, limiting their ability to support workflows in which users progressively refine a track. We introduce AURA, a unified multimodal framework for conversational music editing. AURA uses a multimodal large language model to interpret the complete dialogue history, an optional image, and reference audio, distilling the editing intent into compact concept tokens. A concept-to-audio module injects these tokens and frame-aligned reference features into a frozen MusicGen backbone, enabling precise edits while preserving unaf...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.14344v1 · Indexed about 1 hour ago