ChengRang

Agent Lightning

AI Search & Research Open Source

Microsoft Research Asia open-sourced this rebuilt agentic reinforcement learning framework on October 7, 2026, introducing the Harnessed Agentic RL paradigm in which the same agent harness used at deployment takes part directly in training, so the agent never has to be reimplemented inside the training framework. It places an OpenAI compatible LLM proxy between the agent and the model, leaving existing harness code untouched and usually requiring nothing more than pointing the model endpoint at the proxy. The whole framework is about 3,500 lines, made of an API gateway, a rollout controller and a customized trainer built on verl. Agents run as standard Kubernetes jobs with no dependency on paid sandbox services, and Collocated Async RL lets rollout and model updates share one pool of GPUs for roughly twice the end-to-end speed of synchronous RL. An end-to-end run on SWE-smith, mini-SWE-agent and Qwen3.5-9B lifted Pass@1 on SWE-bench Verified from 41.8 percent to 56.4 percent using only about 6,000 training samples

Reinforcement LearningAgent TrainingHarnessed Agentic RLKubernetesOpen Source Framework
Visit Agent Lightning

Disclaimer: Review content represents our editorial team's views and experience, not commercial recommendation or investment advice. Product info and pricing may change; refer to official sources.

Overview

Agent Lightning v1.0 is the agentic reinforcement learning framework Microsoft Research Asia open-sourced on October 7, 2026, rebuilt from the ground up in roughly 3,500 lines of code. It introduces and formally defines a training paradigm called Harnessed Agentic RL, in which whichever agent harness you use at deployment is the harness that takes part directly in reinforcement learning, removing the need to reimplement the agent inside the training framework.

To see why that matters, look at the assumption traditional agentic RL rests on. Early systems such as verl, AReaL and slime were built on the premise that the training framework owns the interaction loop with the environment. In a ReAct-style loop the model generates an action, the environment returns an observation, the observation is appended to the context and the model generates the next one, so a whole rollout maps onto one continuous token trajectory.

Real harnesses have outgrown that assumption. Coding agents such as mini-SWE-agent, OpenHands, OpenCode, Claude Code and Codex each bring their own context management, tool protocols, execution logic and dependencies, as do general-purpose agent systems. Rebuilding one for training is expensive, and the rebuilt agent may not behave the same way as the one that gets deployed. Agent Lightning takes a different route: place an LLM proxy between the agent and the model, let the agent run exactly as before, and simply point the endpoint that used to call the model API at that proxy, so the training framework can observe and record its model calls. v1.0 formalizes this as Harnessed Agentic RL.

Key Features

Use Cases

Pros

Pricing

Provided free on GitHub as an open-source project, usable, modifiable and self-hostable. Agent execution relies on self-managed Kubernetes clusters, cloud Kubernetes or local infrastructure, whose compute bills under its own rates; the framework itself depends on no paid sandbox service.

Summary

What Agent Lightning v1.0 targets is a structural mismatch in agentic RL: the training framework is bound to the environment loop, while production harnesses stopped letting the training framework touch that loop long ago. Its fix is restrained to the point of surprise, no rewrite and no takeover, just a proxy in the middle that lets the training side observe calls that already exist.

The engineering choices around it are equally practical. 3,500 lines keep the framework at a readable scale, Collocated Async RL resolves the dilemma between idle GPUs and separate GPU pools, and running agents as standard Kubernetes jobs removes the commercial sandbox bill that is easiest to lose control of when rollouts scale. The official example gaining 14.6 points from 6,000 samples also shows the approach works on modest data.

The fit is fairly specific: teams that already run a production-grade agent, want to keep improving it with reinforcement learning, and have no intention of rewriting it for training.

Version History

Category
AI Search & Research
Pricing
Open Source
Tags
Reinforcement Learning · Agent Training · Harnessed Agentic RL
Website

Related Tools