AI agents that optimize their own working environments tend to overspecialize on their test tasks. A new method from Google Cloud AI Research and several universities aims to prevent that issue while cutting compute costs.
Modern AI agents use a harness to control what the model sees at each step. Until recently, humans reviewed failed runs and patched these harnesses manually. Newer automated methods use language models to rewrite the harness repeatedly based on test feedback.
This recursive self-improvement causes agents to memorize their training tasks. Their scores on training tasks rise while gains on unseen tasks shrink or disappear. The search often memorizes patterns fitting only one benchmark and adds unnecessary complexity.
Read nextGoogle Restricts Gemini Models Across Free and Cheaper Tiers Starting October 2026A new approach limits harness edits
Researchers tested a new approach called Regularized Recursive Self-Improvement of Agent Harnesses. RRSI works on both ends of the optimization loop while leaving the harness fully editable. A budget caps the number of independent edits a candidate can bundle at once and shrinks over time.
A critic reviews every proposal and discards any that hardcode task names or solutions. Another rule accepts higher compute costs only if they bring measurable performance gains. Components that no longer help are removed from the system.

Tests showed performance gains on unseen tasks
The team tested RRSI across eight benchmarks using a frozen Claude Opus 4.8 model. RRSI gained up to 14.1 points on trained tasks and up to 4.7 points on five unseen benchmarks. It also used about 30 percent fewer tokens at runtime than the unregularized version.
Other optimization methods performed well on training tasks but failed on new ones. RRSI posted the smallest training gain of all variants and landed well above the baseline on unseen tasks. A coding harness optimized with Gemini 3.5 Flash also raised the accuracy of Gemini 3.1 Flash Lite from 11.2 to 14.6 points.
The study only covers harnesses built around frozen models and excludes changing model weights. The research code is now available on GitHub.



