Overview
Agent Lightning v1.0 is the agentic reinforcement learning framework Microsoft Research Asia open-sourced on October 7, 2026, rebuilt from the ground up in roughly 3,500 lines of code. It introduces and formally defines a training paradigm called Harnessed Agentic RL, in which whichever agent harness you use at deployment is the harness that takes part directly in reinforcement learning, removing the need to reimplement the agent inside the training framework.
To see why that matters, look at the assumption traditional agentic RL rests on. Early systems such as verl, AReaL and slime were built on the premise that the training framework owns the interaction loop with the environment. In a ReAct-style loop the model generates an action, the environment returns an observation, the observation is appended to the context and the model generates the next one, so a whole rollout maps onto one continuous token trajectory.
Real harnesses have outgrown that assumption. Coding agents such as mini-SWE-agent, OpenHands, OpenCode, Claude Code and Codex each bring their own context management, tool protocols, execution logic and dependencies, as do general-purpose agent systems. Rebuilding one for training is expensive, and the rebuilt agent may not behave the same way as the one that gets deployed. Agent Lightning takes a different route: place an LLM proxy between the agent and the model, let the agent run exactly as before, and simply point the endpoint that used to call the model API at that proxy, so the training framework can observe and record its model calls. v1.0 formalizes this as Harnessed Agentic RL.
Key Features
- The Harnessed Agentic RL paradigm: An OpenAI compatible LLM proxy sits between the agent and the model, leaving existing harness code unchanged while the training framework observes and records the prompts, responses and log probabilities behind each model call. The environment interaction loop stays with the harness
- A complete control plane in about 3,500 lines: Three core components make up the framework: an API gateway that stores rollouts, models and events and serves as the LLM proxy, a rollout controller that starts and manages agent execution as local processes or standard Kubernetes jobs and keeps agent execution separate from the trainer, and a customized trainer built on verl that creates rollouts, collects samples and assembles final training samples through a sample adapter
- Collocated Async RL: Synchronous RL waits for the slowest agent in a batch and leaves GPUs idle, while fully asynchronous RL raises utilization but needs separate GPU pools for rollout and training. v1.0 lets rollout and model updates share one set of GPUs: once enough rollouts are collected the gateway pauses new requests, waits for in-flight ones to finish, and rollout resumes after the update. The transition is transparent to the external harness, and experiments show roughly twice the end-to-end speed of synchronous RL while using fewer GPUs than conventional asynchronous RL
- Native Kubernetes job scheduling: Agents run directly as standard Kubernetes jobs, reusing self-managed clusters, cloud Kubernetes or local infrastructure instead of commercial sandbox services such as Modal Sandbox or E2B. Cost does not climb out of control as rollout scale grows, and the whole pipeline stays open source and reproducible
- Answers for four real-harness problems: Because the training system only observes a series of LLM request and response pairs, one rollout can split into a variable number of training samples, creating four challenges: retokenization and sample merging, advantage calculation, loss normalization and training backend scheduling. v1.0 addresses each, with rollout-level advantage combined with rollout-level normalization achieving a higher validation reward and more stable policy entropy in experiments
- End-to-end example with a measured gain: A full pipeline on SWE-smith, mini-SWE-agent and Qwen3.5-9B covers data cleaning, environment construction, reward-hacking safeguards and RL training. With roughly 6,000 training samples and no need for large-scale compute, RL training alone lifted Pass@1 on SWE-bench Verified from 41.8 percent to 56.4 percent, a gain of 14.6 percentage points
Use Cases
- [object Object]
- [object Object]
- [object Object]
- [object Object]
- [object Object]
Pros
- Open source, with the whole framework at roughly 3,500 lines, small enough to read, modify and extend
- Harnessed Agentic RL lets the deployment harness take part in training directly, removing the engineering cost of rebuilding the agent inside a training framework
- Low integration cost, since for an existing harness pointing the model endpoint at the proxy is usually enough
- Agents run as standard Kubernetes jobs, reusing existing clusters or local infrastructure with no dependency on paid sandbox services
- Collocated Async RL shares GPUs between rollout and updates, roughly doubling end-to-end speed versus synchronous RL while using fewer cards
- Explicit handling for all four real-harness problems: retokenization, advantage calculation, loss normalization and backend scheduling
- The end-to-end example gained 14.6 percentage points on SWE-bench Verified from about 6,000 samples, a notable data efficiency result
- The trainer builds on verl, so teams familiar with that ecosystem get up to speed faster
Pricing
Provided free on GitHub as an open-source project, usable, modifiable and self-hostable. Agent execution relies on self-managed Kubernetes clusters, cloud Kubernetes or local infrastructure, whose compute bills under its own rates; the framework itself depends on no paid sandbox service.
Summary
What Agent Lightning v1.0 targets is a structural mismatch in agentic RL: the training framework is bound to the environment loop, while production harnesses stopped letting the training framework touch that loop long ago. Its fix is restrained to the point of surprise, no rewrite and no takeover, just a proxy in the middle that lets the training side observe calls that already exist.
The engineering choices around it are equally practical. 3,500 lines keep the framework at a readable scale, Collocated Async RL resolves the dilemma between idle GPUs and separate GPU pools, and running agents as standard Kubernetes jobs removes the commercial sandbox bill that is easiest to lose control of when rollouts scale. The official example gaining 14.6 points from 6,000 samples also shows the approach works on modest data.
The fit is fairly specific: teams that already run a production-grade agent, want to keep improving it with reinforcement learning, and have no intention of rewriting it for training.
Version History
- Agent Lightning v1.0 open-sourced (2026-10-07): Microsoft Research Asia introduced the Harnessed Agentic RL training paradigm and open-sourced the fully rebuilt v1.0, letting the same agent harness used at deployment take part directly in reinforcement learning. At roughly 3,500 lines, the framework consists of an API gateway, a rollout controller and a customized trainer built on verl, with an OpenAI compatible LLM proxy inserted between agent and model so existing harness code needs no changes. Collocated Async RL lets rollout and model updates share one GPU pool for roughly twice the end-to-end speed of synchronous RL, and agents run as standard Kubernetes jobs with no dependency on commercial sandbox services. It addresses four real-harness challenges: retokenization and sample merging, advantage calculation, loss normalization and backend scheduling. An end-to-end run on SWE-smith, mini-SWE-agent and Qwen3.5-9B lifted Pass@1 on SWE-bench Verified from 41.8 percent to 56.4 percent using about 6,000 training samples