How LLMs Are Trained After Pretraining: SFT, Reward Models, and RL Without the Alphabet Soup
A practical map of LLM post-training: how SFT, reward models, RL (PPO, GRPO), DPO, and RLVR fit together, and why a reward model is not RL.
We haven't written up this one. HackerNoon has the full story — the link below goes straight to it.