Fine-tune a search agent with multi-turn RL on Amazon SageMaker AI
Fine-tuning teaches a small search agent your tools and environment, giving it the reliability of a frontier model at lower latency and cost. In this post, we fine-tune an LLM-powered search agent with multi-turn reinforcement learning (MTRL) on Amazon SageMaker AI and share the gains we measured in retrieval quality and reliability.
Amazon SageMaker AI offers a solution for fine-tuning large language models (LLMs) to create search agents capable of multi-turn interactions. These agents can autonomously decide on retrieval strategies, refine their approach based on previous retrievals, and optimize decisions across multiple turns. However, traditional methods like supervised fine-tuning (SFT) and single-turn reinforcement learning (RL) fail to capture the interdependent decisions made by search agents.
Multi-turn reinforcement learning (MTRL) is the ideal training approach, optimizing the agent's performance over a full multi-turn trajectory while considering environment-specific behavior and output quality. In this post, Amazon SageMaker AI MTRL is used to fine-tune a Qwen3.6-27B model for an enterprise search setting. The search agent relies on two tools: lexical search (BM25) for keyword-based queries and vector search for semantic queries.
To train the model, the author utilized two datasets: FRAMES for multi-hop factoid QA synthesis and BRIGHT for reasoning-intensive retrieval tasks. The training setup involved configuring datasets, a reward function, and MTRL job settings. The datasets included multi-hop factoid QA requiring synthesis across multiple Wikipedia articles and reasoning-intensive retrieval across 12 domains.
The reward function focused on retrieval quality and was designed to encourage efficient search behavior. The MTRL job configuration allowed for modular agent-environment integration, serverless execution at per-token pricing, and asynchronous rollout with bounded off-policy staleness. By using Amazon SageMaker AI MTRL, the authors were able to create a high-performing, cost-effective search agent with reliable multi-turn behavior.
Written by urgent.news from AWS Machine Learning's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.