How LLMs Are Trained After Pretraining: SFT, Reward Models, and RL Without the Alphabet Soup
A practical map of LLM post-training: how SFT, reward models, RL (PPO, GRPO), DPO, and RLVR fit together, and why a reward model is not RL.
After the initial training phase, known as pretraining, Large Language Models (LLMs) undergo additional steps to improve their performance on specific tasks and ensure alignment with human preferences. These steps include supervised fine-tuning (SFT), reward models, and reinforcement learning without explicit policies (RL without the Alphabet Soup). Here's a breakdown of the process:
1. Pretraining: This is the initial stage where the model learns to generate text by predicting the next token in a sequence. It's essentially a giant autocomplete system. The model doesn't answer questions yet; it merely continues a pattern.
2. Supervised Fine-Tuning (SFT, also known as Instruction Tuning): In this stage, the model is taught to answer questions instead of merely continuing a sentence. The input data is structured in instruction-answer format, and the model learns to generate appropriate responses. This step is crucial as it lays the foundation for the model to understand and answer questions effectively.
3. Preference Alignment (RLHF - Reinforcement Learning from Human Feedback): This stage aims to teach the model what people like and dislike. The process involves collecting user feedback (positive and negative preferences) on pairs of answers, training a reward model, and then refining the model's policy. This approach differs from traditional reinforcement learning (RL) because it doesn't involve reward functions or rollouts. Instead, it's a supervised learning task where the model learns to score answers.
4. RLVR (Reward Model-based Value Function Refinement): This step focuses on improving the model's performance on tasks where the answer can be automatically verified, such as math or code. The reward for these answers comes from a checker rather than a trained model. This stage is an alternative to the third step and provides a different way to achieve the same preference-alignment goal.
In summary, the process of training LLMs after pretraining involves several distinct stages: SFT for answering questions, preference alignment (RLHF) for aligning the model with human preferences, RLVR for improving performance on verifiable tasks, and DPO (Direct Preference Optimization) as an alternative method for preference alignment. Each stage builds upon the previous one, ultimately leading to a more effective and user-aligned language model.
Written by urgent.news from HackerNoon's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.