No AI summary available for this article.
Why It Matters
When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs.
Provenance
Discovered via ArXiv and published by ArXiv.
Key Claims
Original description
When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. The total consumption of a task is therefore hard to predict before execution and the prediction must be revised as the run unfolds. In this paper, we propose TokenCast, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it...
Discovered via ArXiv
Research papers and preprints from arXiv.
Publisher: arxiv.org
ID: http://arxiv.org/abs/2609.35760v1 · Indexed 40 minutes ago