ExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds
ExplorationBench evaluates AI systems' scientific exploration in two verifiable 'Alien Worlds' sandboxes whose rules conflict with pre-training, so recall alone cannot solve tasks; of 10 systems tested, the strongest learn unfamiliar rules but gains can stall…
ExplorationBench is a new benchmark measuring whether AI systems can discover new knowledge through exploration rather than recalling pretraining. It uses two executable 'Alien Worlds' sandboxes whose rules conflict with familiar knowledge, so hypotheses can be verified exactly in the environment and recall from pretraining is insufficient. The benchmark comprises AlienCode (31 discovery targets, 70 tasks) and AlienLogic (24 discovery targets, 70 tasks), totaling 140 tasks over 55 discovery targets. Each sandbox supplies a deliberately flawed manual, environmental/task-specific feedback, and a tool-call schema, forcing systems to explore before solving held-out tasks. Ten AI systems were evaluated: the strongest can acquire and apply unfamiliar rules, but performance varies substantially across trajectories, and continued exploration can stall or even reverse earlier gains. Both sources (Hugging Face daily papers, 2026-09-23; arXiv cs.AI/cs.LG/cs.CL, 2026-09-24) report identical figures and findings with no discrepancies.
- Benchmark name: ExplorationBench, measuring AI systems' exploration in verifiable 'Alien Worlds'.
- Two sandboxes: AlienCode with 31 discovery targets and 70 tasks; AlienLogic with 24 discovery targets and 70 tasks.
- Totals: 140 tasks across 55 discovery targets, as stated in the Hugging Face report.
- Each sandbox includes a flawed manual, environmental/task-specific feedback, and a tool-call schema; systems then solve held-out tasks.
- Rules conflict with pre-training knowledge, so tasks cannot be solved by recall alone.
- 10 AI systems were evaluated; the strongest can acquire and apply unfamiliar rules.
- Performance varies substantially across trajectories, and continued exploration can stall or reverse earlier gains.
- Sources agree on all figures and findings; no discrepancies between the 2026-09-23 Hugging Face report and the 2026-09-24 arXiv report.
Coverage timelineoldest first · each row is one article
- · 3d agoExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds
Hugging Face daily papers· 48
ExplorationBench evaluates AI systems' scientific exploration in verifiable 'Alien Worlds' sandboxes where recall cannot solve tasks; top systems still stall or reverse gains.
- · 2d agoExplorationBench: Measuring AI Systems' Exploration in Verifiable Alien Worlds
arXiv cs.AI / cs.LG / cs.CL· 52
ExplorationBench tests whether AI systems discover new rules in verifiable alien worlds instead of recalling pretraining.