Microsoft Research Asia has open-sourced Agent Lightning v1.0, a reinforcement learning framework designed to train AI agents with the same software harnesses they use in deployment. The company says developers can connect an existing agent by pointing its model endpoint at Agent Lightning’s proxy, without rebuilding the agent inside the training framework.
An agent harness manages tasks such as tool use, execution and context. Agent Lightning’s proxy records the agent’s model calls for training while the harness continues to run. Microsoft calls this approach Harnessed Agentic RL.
The framework has about 3,500 lines of code
The rebuilt framework is about 3,500 lines of code, according to Microsoft. Its API gateway records model calls and training data, a rollout controller manages agent runs, and a trainer assembles samples and updates the model. Agents can run as local processes or standard Kubernetes jobs on self-managed clusters, cloud Kubernetes or local infrastructure.
Microsoft also introduced a method that lets agent runs and model updates share the same GPUs. When enough runs have finished, the gateway pauses new model requests, lets requests in progress finish and resumes agent runs after the update. Microsoft says its experiments showed about twice the end-to-end speed of synchronous reinforcement learning while using fewer GPUs than conventional asynchronous reinforcement learning.

Reinforcement learning improved coding performance
In a coding-agent example using SWE-smith, mini-SWE-agent and Qwen3.5-9B, Microsoft says reinforcement learning raised the model’s Pass@1 score on SWE-bench Verified from 41.8% to 56.4%. That is a 14.6 percentage point gain from a training set of about 6,000 samples.
Agent Lightning v1.0 is available as open-source software.



