Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses
Training AI agents with reinforcement learning can be challenging because their tools, context, and decision-making are managed by complex frameworks. Agent Lightning connects existing agents to RL training, making it easier to improve them without rebuilding them. The post Agent Lightning v1.0: A 3,500-Line Lightweight Agentic RL Framework for Training Agents with Real Harnesses appeared first…
Microsoft Research Asia has unveiled a lightweight, fully rebuilt Agent Lightning v1.0 framework for harness-based agentic RL. This new training paradigm integrates the same agent harness directly into reinforcement learning, eliminating the need to reimplement agents inside the training framework. Agent Lightning v1.0 boasts a compact codebase of roughly 3,500 lines and offers native Kubernetes support, enabling agents to run as standard Kubernetes jobs on various infrastructure.
The framework achieved a notable 14.6 percentage point improvement in SWE-bench Verified performance for Qwen3.5-9B, raising Qwen3.5-9B from 41.8% to 56.4% Pass@1 using just 6,000 training samples. Traditional agentic RL systems, which embed the interaction loop within the training framework, face challenges in adapting to real harnesses, such as retokenization, advantage calculation, loss normalization, and backend scheduling.
Harnessed Agentic RL addresses these challenges by allowing the harness used in deployment to participate directly in reinforcement learning during training.
Written by urgent.news from Microsoft Research's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.