No AI summary available for this article.
Why It Matters
One deployed language model must often serve many compute budgets, yet serving each budget still means a separate training or compression run per point.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
One deployed language model must often serve many compute budgets, yet serving each budget still means a separate training or compression run per point. We train a Telescopic Language Model (TLM) to be that continuum: a nested-capacity Transformer supervised by stochastic prefix supervision with a full anchor. At every step, one randomly truncated prefix of the capacity axis is trained against the full next-token target, alongside one full-capacity pass, so the trained artifact is a valid language model at every depth. Two forward-backward passes per step, no architectural change, nothing extr...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.35769v1 · Indexed 41 minutes ago