CodeScene AI Agents Refactor 300K Lines of Code in Three Weeks
Tech

CodeScene AI Agents Refactor 300K Lines of Code in Three Weeks

TechNews Editorial
TechNews EditorialSep 30, 2026 · 2 min read
Share

Why it matters

The case study shows that agents can refactor massive codebases using deterministic harnesses, though experts debate whether the results apply to standard production systems.

The facts

  • Coding agents refactored a 300,000-line C codebase over three weeks for roughly $4,000.
  • The process raised the Code Health score from 5.6 to 10.0 across 2,903 commits.
  • A Lund University study will use the resulting code versions to compare student feature implementation.

CodeScene published a case study detailing how coding agents refactored a 300,000-line C codebase over three weeks. The project cost roughly $4,000 in tokens. It produced 2,903 commits across 726 files and modified 252,055 lines. The codebase shifted from a Code Health score of 5.6 to 10.0. The target was Street Fighter III: 3rd Strike, using an open-source decompilation.

Adam Tornhill, CodeScene's founder, stated this was his first time seeing superhuman AI performance at scale in three decades. Two main mechanisms drove the work. First, the CodeHealth MCP Server provided a deterministic score to optimize. Second, a replay-trace harness compared rollback state hashes frame by frame to check behavior after every change.

Agents build new refactoring recipes

Agents built an emergent refactoring playbook during the process. They ended with 22 recipes and 82 supporting notes. The recipes included Shared Index Range, Action Parameter, and Uniform Step Table. Failed attempts were also recorded by the system. Model choice impacted the results significantly. The team used Claude Opus for the bulk of the work, noting that smaller models plateaued and could not move past local optimums.

Practitioners debate scope and testing

Practitioners on LinkedIn reacted with a sharp split over what the study proves. Mats Iremark called the combination almost like cheating. Skeptics questioned merge status and whether the work applies to production code earning money. Daniel Webb, one of the two engineers on the project, confirmed the work merged to main on a fork through 54 pull requests. Other critics questioned the recipes, methodology, and unmeasured architecture.

A game replay test advances original and refactored versions together, showing matching character positions and identical scenery at the same instant.
Illustration: AI & Tech News

The authors raised their own questions regarding non-functional behavior such as framerate, memory use, and input latency. Webb noted that a performance specialist is now joining the effort. The replay-trace harness succeeded because a decompiled game allows deterministic frame-by-frame replay. Most legacy systems lack such an oracle, which normally makes refactoring risky.

The team bypassed a Gov.UK marine licensing codebase because it was too healthy. They chose the game partly because they play it. The resulting two versions of the system, one at Code Health 5.6 and one at 10.0, will be used in a study with Lund University where students implement features using frontier models.

Newsletter

Get the best AI & tech news daily

A concise daily digest. Unsubscribe anytime.

We use your email only to send this newsletter.

Keep reading