Agent in a Bottle: Can LLM Agents Turn Their Capabilities Into Cheap, Scalable Artifacts?
BOTTLED benchmarks whether LLM agents can convert large repetitive workloads into cheaper reusable solutions.
BOTTLED evaluates whether LLM agents can turn a large unlabelled workload into a cheaper reusable solution under fixed time, compute, and API budgets. Across ten models and three tasks, strong zero-shot scores do not reliably predict this ability: 48 of 60 runs score below the lower bound of their model's zero-shot 95% confidence interval, and 31 of 60 underperform the stronger of two small-model distillation baselines. On query-product relevance classification, Opus 5 retains about 82% of its zero-shot macro-F1 at roughly 657 times lower reported cost. It also recovers about 94% of Jev's macro-F1 at a quarter of Jev's projected full-workload cost.
- Agents must complete an unlabelled workload under time, compute, and API budgets.
- 48 of 60 runs fall below their model's zero-shot confidence interval.
- 31 of 60 runs lose to the stronger small-model distillation baseline.
- Opus 5 keeps about 82% macro-F1 at roughly 657 times lower cost.
Full article247 words · extracted from arxiv.org · click to collapse
Large language models (LLMs) can solve many narrow tasks, but querying them separately for millions of related instances can be prohibitively expensive. Can LLM agents autonomously create cheaper solutions for such workloads? We call this ability "bottling": the ability to turn general capabilities into task-specific solutions that balance answer quality and amortised cost. We introduce BOTTLED, a benchmark in which agents receive an entire unlabelled workload and must complete it under fixed time, compute and LLM API budgets. Agents choose their own approach, such as training a small model or writing a reusable program. Across ten models and three tasks, we find that strong zero-shot task performance does not reliably translate into strong bottling capabilities. Models with similar zero-shot scores can differ substantially after bottling, and 48 of 60 bottling runs score below the lower bound of the 95% confidence interval of their model's zero-shot performance. Moreover, 31 of 60 runs underperform the stronger of two small-model distillation baselines with the same token budget. Nevertheless, bottling can yield substantial savings: on query-product relevance classification, Opus 5 retains about 82% of its zero-shot macro-F1 at roughly 657 times lower reported cost. Bottling is also competitive with Jev, a "system one" model built especially for cheap, repetitive inference: Opus 5 on the same task recovers about 94% of Jev's macro-F1 at a quarter of Jev's projected full-workload cost. BOTTLED provides a basis for evaluating and improving agents' ability to invest limited resources in reusable solutions for large, repetitive workloads.
Text extracted automatically; images, tables and formatting may be missing. Original: https://arxiv.org/abs/2610.08775