Keep the agent, change the training path
Many reinforcement-learning systems expect the training framework to control an agent’s interaction with its environment. That works for a simple loop: the model chooses an action, receives an observation, then chooses again. But real agents bring their own context management, tool protocols, execution logic and dependencies.
Rebuilding those systems inside a training framework takes work and can produce an agent that behaves differently from the one used in production. Agent Lightning’s alternative is to route the agent’s model requests through an LLM proxy. The agent keeps its existing code; the framework observes and records its requests and responses.
The v1.0 architecture has three components:
Figure 1. Traditional agentic RL compared with Harnessed Agentic RL. In traditional agentic RL, the training framework manages the environment and the agent loop. In Harnessed Agentic RL, the harness manages both.
Source: microsoft.com
Using the real pipeline leaves the trainer with an awkward view of the data: it sees LLM requests and responses, not the whole agent run. One run may yield different numbers of training samples, creating several problems:
One GPU pool, asynchronous runs
Agent runs take different amounts of time. In synchronous training, the system waits for the slowest agent in a group. Fully asynchronous training can improve utilization but typically needs separate GPU pools for agent runs and training.
Agent Lightning v1.0 allows both to run asynchronously on one set of GPUs. Once enough runs are collected, the API Gateway pauses new requests and waits for in-flight ones to finish while the model is updated. Agent execution then resumes. The researchers report roughly twice the end-to-end speed of synchronous training, using fewer GPUs than conventional asynchronous training.
Figure 2. The Agent Lightning v1.0 system architecture showing the API gateway, rollout controller, and customized trainer.
Source: microsoft.com
The framework also runs agents as standard Kubernetes jobs rather than relying on commercial sandboxes such as Modal Sandbox or E2B. Researchers can use their own clusters, cloud Kubernetes or local infrastructure.
For a coding-agent test, the researchers combined SWE-smith, mini-SWE-agent and Qwen3.5-9B in a pipeline covering data cleanup, runtime setup, reward-hacking protections and reinforcement learning. About 6,000 training samples raised the model’s SWE-bench Verified score from 41.8% to 56.4% Pass@1—a gain of 14.6 percentage points.
Figure 3. Synchronous RL, asynchronous RL, and Collocated Async RL compared. Collocated Async RL raises utilization while occupying fewer GPUs.
Source: microsoft.com
The harder question is what the score measures
I think the most consequential claim here is not the score increase; it is that reinforcement learning can use an agent’s actual pipeline without making that pipeline part of the trainer. If that holds beyond this example, it addresses a practical mismatch between how agents are trained and how they are run.
The paper also reports that calculating advantages and normalizing losses at the run level produced higher validation rewards than doing so at the sample level, while helping keep policy entropy stable. Those choices matter because the pipeline itself determines how many samples a run produces.
What I’d want to know is how robust the result is across agents and tasks. The report gives one coding-agent example, and the framework’s design shifts complexity into handling variable runs, tokenization and scheduling. A compact implementation is appealing, but the real test is whether teams can keep their own pipelines intact without those edge cases becoming the new integration burden.
Figure 4. The Rollout Controller in Agent Lightning v1.0 provides native Kubernetes support, running agents directly as standard Kubernetes jobs.
Source: microsoft.com
The work is described in Agent Lightning v1.0: Towards Harnessed Agentic RL. The listed authors are Zhiyuan He, a Software Engineer II, and Yuqing Yang.
Figure 5. Pass rate and policy entropy for Qwen3.5-9B on the SWE-smith validation set.
Source: microsoft.com
Daily AI news
Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.
Only what matters — every day
Follow on X