No AI summary available for this article.
Why It Matters
Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
Mixture-of-Experts (MoE), popularized by large language models, is a promising paradigm for scaling visual generative models. However, conventional token-wise MoE routes tokens independently within a homogeneous expert pool and regularizes expert usage toward uniformity, making it poorly matched to video data that is spatiotemporally redundant and semantically long-tailed. We show that existing visual MoEs fall into a uniformity trap: semantically under-organized routing, compounded by uniform expert-usage regularization, scatters coherent patches across disparate experts, causing routing frag...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.38140v1 · Indexed about 1 hour ago