Urgent.News

What's breaking now, across thousands of outlets.

AI

How LLMs Are Trained After Pretraining: SFT, Reward Models, and RL Without the Alphabet Soup

A practical map of LLM post-training: how SFT, reward models, RL (PPO, GRPO), DPO, and RLVR fit together, and why a reward model is not RL.

How LLMs Are Trained After Pretraining: SFT, Reward Models, and RL Without the Alphabet Soup

After the initial training phase, known as pretraining, Large Language Models (LLMs) undergo additional steps to improve their performance on specific tasks and ensure alignment with human preferences. These steps include supervised fine-tuning (SFT), reward models, and reinforcement learning without explicit policies (RL without the Alphabet Soup). Here's a breakdown of the process:

1. Pretraining: This is the initial stage where the model learns to generate text by predicting the next token in a sequence. It's essentially a giant autocomplete system. The model doesn't answer questions yet; it merely continues a pattern.

2. Supervised Fine-Tuning (SFT, also known as Instruction Tuning): In this stage, the model is taught to answer questions instead of merely continuing a sentence. The input data is structured in instruction-answer format, and the model learns to generate appropriate responses. This step is crucial as it lays the foundation for the model to understand and answer questions effectively.

3. Preference Alignment (RLHF - Reinforcement Learning from Human Feedback): This stage aims to teach the model what people like and dislike. The process involves collecting user feedback (positive and negative preferences) on pairs of answers, training a reward model, and then refining the model's policy. This approach differs from traditional reinforcement learning (RL) because it doesn't involve reward functions or rollouts. Instead, it's a supervised learning task where the model learns to score answers.

4. RLVR (Reward Model-based Value Function Refinement): This step focuses on improving the model's performance on tasks where the answer can be automatically verified, such as math or code. The reward for these answers comes from a checker rather than a trained model. This stage is an alternative to the third step and provides a different way to achieve the same preference-alignment goal.

In summary, the process of training LLMs after pretraining involves several distinct stages: SFT for answering questions, preference alignment (RLHF) for aligning the model with human preferences, RLVR for improving performance on verifiable tasks, and DPO (Direct Preference Optimization) as an alternative method for preference alignment. Each stage builds upon the previous one, ultimately leading to a more effective and user-aligned language model.

Written by urgent.news from HackerNoon's reporting — not their text. Machine-written — may contain errors; check the original before relying on it.

Read the original at hackernoon.com →

More in AI

AI in schools brings back pen and paper

A child sitting down to complete homework in 2026 has more assistance within reach than any previous generation. A chatbot can explain a maths problem, suggest an essay structure, summarise a chapter and correct a paragraph within seconds.

A recent controversy surrounding an artificial intelligence (AI) system's grading for an English essay test has sparked concerns over fairness and accuracy in automated evaluations. The issue came to light when a student took the College Scholastic Ability Test (CSAT) mock exam, specifically the English section, and received an unexpectedly high score for an essay response that was later found to contain several instances of specific phrases or "buzzwords" that are often favored in evaluations. The student expressed frustration upon learning that despite their teacher grading the same essay as a C, the AI system awarded it an A. This discrepancy has raised questions about the reliability of AI in assessing student performance, particularly when it appears to prioritize certain keywords over the logical coherence or overall quality of the response. Further investigation revealed that the AI system seemed to give more weight to the presence of certain "buzzwords" rather than the content or logical structure of the essay. This approach has been criticized for potentially rewarding students who are adept at identifying and incorporating favored phrases into their work, rather than those who demonstrate a deeper understanding or more nuanced expression of the subject matter. The Korea Institute of Curriculum and Evaluation (KICE), the organization responsible for developing and grading the CSAT, has acknowledged the issue and is taking steps to address concerns over the AI grading system. KICE officials have stated that they are working to refine the AI evaluation algorithm to ensure that it assesses responses based on a more balanced and comprehensive set of criteria, rather than relying heavily on specific keywords. The incident highlights the challenges and limitations of using AI in educational assessments, particularly in evaluating subjective or creative tasks such as essay writing. As educators and policymakers continue to integrate AI into various aspects of education, ensuring the fairness, accuracy, and reliability of AI systems in grading and evaluation will be crucial. The goal is to create a system that supports and fairly assesses student learning, without inadvertently encouraging superficial strategies or disadvantaging students who approach problems differently.

More from Friday 14 August →