No AI summary available for this article.
Why It Matters
We distill Kokoro-82M, a widely used open text-to-speech model with 54 voices, into Paradee, an 8.07M-parameter model that speaks one of them.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
We distill Kokoro-82M, a widely used open text-to-speech model with 54 voices, into Paradee, an 8.07M-parameter model that speaks one of them. Paradee keeps Kokoro's architecture with much narrower layers, and each of its two halves is trained separately against the frozen teacher. It has 10x fewer parameters and needs 15x less compute. We first synthesize a corpus with the teacher and keep its durations, pitch, energy and phoneme features. We then train a small text side to predict these values, and a small decoder to turn the teacher's saved values into the teacher's audio, first with spectr...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2610.06817v1 · Indexed about 1 hour ago