A Closer Look at Agentic BBO: Benchmarking LLM Agents for Black-Box Optimization
AgenticBBO-Bench compares LLM agents on black-box optimization across five domains, with GPT-6 Astra on the cost-performance frontier.
The paper introduces AgenticBBO-Bench, a unified finite-budget benchmark for LLM agents doing black-box optimization across synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design. Agentic systems scored higher on family averages than direct LLM methods in all five domains and beat the best numerical optimizers in four. Ablations found extra numerical tools did not consistently help, while task semantics were more reliable than specific priors. On a five-task frontier challenge with the Codex harness, GPT-6 Astra and DeepSeek-V4.1-Flash were on the performance-cost Pareto frontier among seven LLMs.
- Benchmark spans synthetic functions, hyperparameter optimization, databases, chips, and molecules.
- Agentic BBO beat direct LLM methods on family-averaged scores in every domain.
- Agents also beat the best numerical optimizers in four of the five domains.
- More tools did not reliably help; task semantics helped more than specific priors.
- Under Codex, GPT-6 Astra and DeepSeek-V4.1-Flash lead the cost-performance frontier.
Full article236 words · extracted from huggingface.co · click to collapse
Black-box optimization (BBO) arises in many scientific and engineering problems where objective evaluations are expensive and limited. Recent large language model (LLM) agents offer a new way to approach BBO by combining task semantics, computation, optimization tools, and feedback-driven decision making, showing great potential due to the integration with mathematically rigorous tools. However, existing agentic BBO studies use different task domains and system configurations, making their results difficult to compare and the effects of individual design choices hard to isolate. We therefore introduce AgenticBBO-Bench, a cross-domain benchmark for agentic BBO spanning synthetic functions, hyperparameter optimization, database tuning, chip design, and molecular design under a unified finite-budget evaluation protocol. In our experiments, agentic BBO achieves higher family-averaged scores than direct LLM-based methods in all five domains and outperforms the best numerical optimizers in four. We further study three factors shaping agent performance: optimization tools, task information and prior knowledge, and the role of the LLM during search. Our results show that additional numerical tools do not consistently improve performance, task semantics are broadly useful while more specific priors are less reliable, and numerical optimizers can effectively absorb gains from search trajectories established by the agent. Finally, we introduce a five-task frontier challenge within AgenticBBO-Bench and evaluate seven LLMs under the Codex agent harness, where GPT-6 Astra and DeepSeek-V4.1-Flash lie on the Pareto frontier of performance and cost among the evaluated models. Our code is available at https://github.com/lamda-bbo/agentic-bbo.
Text extracted automatically; images, tables and formatting may be missing. Original: https://huggingface.co/papers/2610.12183