No AI summary available for this article.
Why It Matters
In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
In diffusion transformers, a class label or a text prompt is embedded once, and the same condition is reused at every denoising step. We ask whether predicted embeddings can serve as this condition instead. Next-Embedding Predictive Autoregression (NEPA) trains a Transformer to predict the next continuous embedding in a sequence. In generation, the clean image follows the noisy image, so its embeddings are the next embeddings after the condition and the noisy image. We train a NEPA model to predict them all at once with Multi-Embedding Prediction, and in Embedding Conditioned Generation, a DiT...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2610.02203v1 · Indexed about 1 hour ago