TokenCast: Forecasting Token Consumption During LLM Agent Execution
TokenCast forecasts LLM agent token use from segment costs, cutting error 14.5% across 96 combinations.
TokenCast learns a composable cost representation for each LLM-agent execution segment, recording that segment's token use and the context growth it adds. Composing segments estimates cumulative cost, including extra input tokens when earlier context is reread by later calls, and the forecast updates from new evidence without additional LLM calls. Mean cumulative prediction time is 32.8 ms per run on SWE-bench Verified. Across four task suites and six agent models, mean absolute error drops 14.5% versus the strongest comparator over 96 combinations, and offline budget control uses 21.3% fewer tokens than a fixed budget at matched completion.
- The same agent task can vary by more than ten times in tokens.
- Forecasts refresh from observed evidence with no extra LLM calls.
- Mean cumulative prediction time is 32.8 ms on SWE-bench Verified.
- Mean absolute error falls 14.5% across 96 model-task combinations.
- Budget replay uses 21.3% fewer tokens than a fixed-budget policy.
Full article200 words · extracted from arxiv.org · click to collapse
When a large language model (LLM) agent executes the same task, token consumption can vary by over an order of magnitude across runs. The agent chooses its next steps based on tool feedback and intermediate results, while the growing context steadily inflates the input size of every subsequent call. The total consumption of a task is therefore hard to predict before execution and the prediction must be revised as the run unfolds. In this paper, we propose TokenCast, which learns a composable cost representation for each execution segment, recording its own consumption and the context growth it introduces. Composing adjacent segments yields a cumulative estimate that captures the extra input cost incurred when context from earlier segments is re-read by every later call. As execution unfolds, newly observed evidence refreshes the forecast, requiring no additional LLM calls and incurring a mean cumulative prediction time of 32.8 ms per run on SWE-bench Verified. Across 4 task suites and 6 agent models, TokenCast's mean absolute error reduction against the strongest comparator averages 14.5% over 96 evaluated combinations. In offline budget-control replay, TokenCast uses 21.3% fewer tokens on average than a fixed-budget policy at matched trace completion. The code is available at https://github.com/DEFENSE-SEU/TokenCast.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2609.35760