CodeScene published a case study detailing how coding agents refactored a 300,000-line C codebase over three weeks. The project cost roughly $4,000 in tokens. It produced 2,903 commits across 726 files and modified 252,055 lines. The codebase shifted from a Code Health score of 5.6 to 10.0. The target was Street Fighter III: 3rd Strike, using an open-source decompilation.
Adam Tornhill, CodeScene's founder, stated this was his first time seeing superhuman AI performance at scale in three decades. Two main mechanisms drove the work. First, the CodeHealth MCP Server provided a deterministic score to optimize. Second, a replay-trace harness compared rollback state hashes frame by frame to check behavior after every change.
Agents build new refactoring recipes
Agents built an emergent refactoring playbook during the process. They ended with 22 recipes and 82 supporting notes. The recipes included Shared Index Range, Action Parameter, and Uniform Step Table. Failed attempts were also recorded by the system. Model choice impacted the results significantly. The team used Claude Opus for the bulk of the work, noting that smaller models plateaued and could not move past local optimums.
Practitioners debate scope and testing
Practitioners on LinkedIn reacted with a sharp split over what the study proves. Mats Iremark called the combination almost like cheating. Skeptics questioned merge status and whether the work applies to production code earning money. Daniel Webb, one of the two engineers on the project, confirmed the work merged to main on a fork through 54 pull requests. Other critics questioned the recipes, methodology, and unmeasured architecture.

The authors raised their own questions regarding non-functional behavior such as framerate, memory use, and input latency. Webb noted that a performance specialist is now joining the effort. The replay-trace harness succeeded because a decompiled game allows deterministic frame-by-frame replay. Most legacy systems lack such an oracle, which normally makes refactoring risky.
The team bypassed a Gov.UK marine licensing codebase because it was too healthy. They chose the game partly because they play it. The resulting two versions of the system, one at Code Health 5.6 and one at 10.0, will be used in a study with Lund University where students implement features using frontier models.



