i
News
News · 2026-10-07

Microsoft’s Agent Lightning trains agents without rebuilding their pipelines

@neuronium_ai @neuronium_ai

Microsoft Research Asia has released Agent Lightning v1.0, an open-source framework for training AI agents with reinforcement learning while keeping their existing pipelines intact. The framework is about 3,500 lines of code and puts a model proxy between the agent and the training system, rather than requiring developers to rebuild the agent inside a training framework. In a coding-agent example, training Qwen3.5-9B on about 6,000 samples raised its SWE-bench Verified Pass@1 score from 41.8% to 56.4%.

Cover: Microsoft’s Agent Lightning trains agents without rebuilding their pipelines

Keep the agent, change the training path

Many reinforcement-learning systems expect the training framework to control an agent’s interaction with its environment. That works for a simple loop: the model chooses an action, receives an observation, then chooses again. But real agents bring their own context management, tool protocols, execution logic and dependencies.

Rebuilding those systems inside a training framework takes work and can produce an agent that behaves differently from the one used in production. Agent Lightning’s alternative is to route the agent’s model requests through an LLM proxy. The agent keeps its existing code; the framework observes and records its requests and responses.

The v1.0 architecture has three components:

API Gateway stores agent runs, models and events, and acts as an OpenAI-compatible LLM proxy. It links model calls to runs and records prompts, responses and log probabilities.
Rollout Controller launches and manages agents as local processes or ordinary Kubernetes jobs, separating agent execution from training.
Customized Trainer, built on verl, collects completed runs and turns them into training data through a sample adapter.
Figure 1. Traditional agentic RL compared with Harnessed Agentic RL. In traditional agentic RL, the training framework manages the environment and the agent loop. In Harnessed Agentic RL, the harness manages both.

Figure 1. Traditional agentic RL compared with Harnessed Agentic RL. In traditional agentic RL, the training framework manages the environment and the agent loop. In Harnessed Agentic RL, the harness manages both.

Source: microsoft.com

Using the real pipeline leaves the trainer with an awkward view of the data: it sees LLM requests and responses, not the whole agent run. One run may yield different numbers of training samples, creating several problems:

Text must be converted back into the token IDs used during execution, but applying a chat template and tokenizer again can change token boundaries.
Advantage calculations at the sample level can overcount runs that produce more samples.
Losses averaged by sample count can give those same runs more weight, even when sample count reflects pipeline behavior rather than training value.
The number and length of samples are known only after a pipeline finishes, complicating scheduling across fixed GPU resources.

One GPU pool, asynchronous runs

Agent runs take different amounts of time. In synchronous training, the system waits for the slowest agent in a group. Fully asynchronous training can improve utilization but typically needs separate GPU pools for agent runs and training.

Agent Lightning v1.0 allows both to run asynchronously on one set of GPUs. Once enough runs are collected, the API Gateway pauses new requests and waits for in-flight ones to finish while the model is updated. Agent execution then resumes. The researchers report roughly twice the end-to-end speed of synchronous training, using fewer GPUs than conventional asynchronous training.

Figure 2. The Agent Lightning v1.0 system architecture showing the API gateway, rollout controller, and customized trainer.

Figure 2. The Agent Lightning v1.0 system architecture showing the API gateway, rollout controller, and customized trainer.

Source: microsoft.com

The framework also runs agents as standard Kubernetes jobs rather than relying on commercial sandboxes such as Modal Sandbox or E2B. Researchers can use their own clusters, cloud Kubernetes or local infrastructure.

For a coding-agent test, the researchers combined SWE-smith, mini-SWE-agent and Qwen3.5-9B in a pipeline covering data cleanup, runtime setup, reward-hacking protections and reinforcement learning. About 6,000 training samples raised the model’s SWE-bench Verified score from 41.8% to 56.4% Pass@1—a gain of 14.6 percentage points.

Figure 3. Synchronous RL, asynchronous RL, and Collocated Async RL compared. Collocated Async RL raises utilization while occupying fewer GPUs.

Figure 3. Synchronous RL, asynchronous RL, and Collocated Async RL compared. Collocated Async RL raises utilization while occupying fewer GPUs.

Source: microsoft.com

The harder question is what the score measures

I think the most consequential claim here is not the score increase; it is that reinforcement learning can use an agent’s actual pipeline without making that pipeline part of the trainer. If that holds beyond this example, it addresses a practical mismatch between how agents are trained and how they are run.

The paper also reports that calculating advantages and normalizing losses at the run level produced higher validation rewards than doing so at the sample level, while helping keep policy entropy stable. Those choices matter because the pipeline itself determines how many samples a run produces.

What I’d want to know is how robust the result is across agents and tasks. The report gives one coding-agent example, and the framework’s design shifts complexity into handling variable runs, tokenization and scheduling. A compact implementation is appealing, but the real test is whether teams can keep their own pipelines intact without those edge cases becoming the new integration burden.

Figure 4. The Rollout Controller in Agent Lightning v1.0 provides native Kubernetes support, running agents directly as standard Kubernetes jobs.

Figure 4. The Rollout Controller in Agent Lightning v1.0 provides native Kubernetes support, running agents directly as standard Kubernetes jobs.

Source: microsoft.com

The work is described in Agent Lightning v1.0: Towards Harnessed Agentic RL. The listed authors are Zhiyuan He, a Software Engineer II, and Yuqing Yang.

Figure 5. Pass rate and policy entropy for Qwen3.5-9B on the SWE-smith validation set.

Figure 5. Pass rate and policy entropy for Qwen3.5-9B on the SWE-smith validation set.

Source: microsoft.com

Daily AI news

Every day we pick what actually matters in AI and explain it plainly — no hype, no filler. Subscribe if you want to follow where the industry is going.

Only what matters — every day

Follow on X