TokenCast Forecasts LLM Agent Token Use From Segment Costs
TokenCast predicts LLM-agent token use from composable segment costs, cutting error 14.5% and using 21.3% fewer tokens than a fixed budget.
TokenCast forecasts token consumption during LLM-agent execution by learning a composable cost for each execution segment, including that segment's own tokens and the context growth it adds. The same agent task can vary by more than an order of magnitude in tokens because tool feedback and growing context change later calls; composing segments estimates cumulative cost, including extra input tokens when earlier context is reread. Forecasts refresh from new evidence without additional LLM calls, and mean cumulative prediction time is 32.8 ms per run on SWE-bench Verified. Across four task suites and six agent models, mean absolute error drops 14.5% versus the strongest comparator over 96 combinations. Offline budget control uses 21.3% fewer tokens than a fixed budget at matched completion. The two reports agree on these figures; one calls the budget result offline replay and the other offline budget control.
- Token use for the same LLM-agent task can vary by more than an order of magnitude (more than ten times) because tool feedback and growing context change later calls.
- TokenCast learns a composable cost for each execution segment, covering that segment's own tokens and the context growth it adds, so cumulative forecasts include extra input tokens when earlier context is reread.
- Forecasts update online from observed evidence with no additional LLM calls.
- Mean cumulative prediction time is 32.8 ms per run on SWE-bench Verified.
- Across 4 task suites and 6 agent models, mean absolute error falls 14.5% versus the strongest comparator over 96 combinations.
- Offline budget replay uses 21.3% fewer tokens than a fixed budget at matched completion.
Coverage timelineoldest first · each row is one article
- · 2d agoTokenCast: Forecasting Token Consumption During LLM Agent Execution
Hugging Face daily papers· 46
TokenCast forecasts LLM agent token use from composable segment costs, cutting error 14.5% on average.
- · 1d agoTokenCast: Forecasting Token Consumption During LLM Agent Execution
arXiv cs.AI / cs.LG / cs.CL· 46
TokenCast forecasts LLM agent token use from segment costs, cutting error 14.5% across 96 combinations.