Google Research Open-Sources RRSI: AI Agents That Improve Their Own Harness Without Overfitting
Google Research open-sourced RRSI, letting frozen-weight agents improve their harness without overfitting benchmarks.
Google Cloud AI Research, with UNC-Chapel Hill, Stanford, and Washington University in St. Louis, open-sourced Regularized Recursive Self-Improvement (RRSI) under Apache 2.0. It lets an LLM agent edit prompts, tools, memory, control flow, and sub-agents while weights stay frozen, using an annealed edit budget, leakage critic, noise floor, cost rule, and pruning. With Claude Opus 4.8, Terminal-Bench 2.1 rose from 74.2% to 80.2% and held-out SWE-bench Verified from 82.0% to 83.8%, with further gains on JobBench, GDPval, and APEX-Agents. Defaults target Claude Opus 4.8 via LiteLLM; the authors call the GitHub release research-grade, not an official Google product.