ML Types Review · Reinforcement Learning · RLHF & RLVR · Intro to NLP
Every ML system learns differently based on what kind of feedback it receives
Learns from labeled examples. You provide input + correct answer. The model learns the mapping.
Email → spam/not spam
Image → cat/dog
Loan application → approve/deny
Finds hidden patterns in data with no labels. The model groups or compresses on its own.
Customer segments
Topic clustering
Anomaly detection
Learns by trial and error in an environment. Gets rewards for good actions, penalties for bad ones.
Game playing (AlphaGo)
Robot locomotion
Ad bidding systems
Most real-world AI systems combine multiple approaches — these are categories, not strict boxes.
Most real-world data is unlabeled — labeling it costs time, money, and human expertise
An agent takes actions in an environment and learns from the reward it receives
What the agent can "see" right now — the current situation.
Chess: current board positions
What the agent decides to do from the available options.
Chess: which piece to move where
The score — positive or negative — the environment gives back.
Chess: +1 for winning, -1 for losing
The reward signal is everything — the agent will optimize for exactly what you measure
"You get what you measure" — the alignment problem starts here.
Every RL agent faces this fundamental tradeoff on every step
Use what you already know to get the best reward you can right now.
You always go to the same restaurant because it's good. But you'll never find a better one.
Try new things to gather information, even at the cost of short-term rewards.
You try a new restaurant each week. Sometimes terrible — but once in a while, life-changing.
RL algorithms use strategies like ε-greedy (explore randomly X% of the time) or decaying exploration (explore a lot early, exploit later as confidence grows) to balance both sides.
Two very different kinds of feedback loops
Supervised: train once, deploy, done. No feedback from the real world after training.
RL: continuous loop — the agent keeps learning from consequences, not from pre-labeled answers.
Any domain where you make sequences of decisions and can define a reward
AlphaGo, AlphaStar (StarCraft II), OpenAI Five (Dota 2). Reward: winning. Discovered strategies no human had ever played.
Teaching robots to walk, grasp, and navigate. Reward: staying upright, reaching targets, picking up objects without dropping.
Netflix, YouTube, Spotify. Reward: watch time, clicks, satisfaction ratings. Each user interaction is feedback for the next suggestion.
Designing new molecule structures. Reward: predicted binding affinity, low toxicity. Explores chemical space faster than human chemists.
Common thread: sequential decisions + measurable outcomes + lots of trial opportunities. If any of these is missing, RL gets much harder.
When you can't write down the reward — you use humans to teach what "good" means
Why RLHF matters for LLMs: You can't write a formula for "helpful, harmless, honest." RLHF lets humans teach those properties directly. GPT-4, Claude, and Gemini all use versions of this process.
The same underlying model behaves very differently once human preferences are trained in
Technically correct — but robotic, generic, and doesn't actually answer the question.
Helpful, practical, conversational, and asks a follow-up. This is what human raters preferred.
The model learned this style from thousands of pairwise human comparisons — not explicit instructions.
What if you could check the answer automatically — no human needed?
For some tasks, correctness is objective and checkable:
These are "verifiable rewards" — no human judgment needed, just a checker.
Model generates a chain-of-thought answer. A verifier checks if the final number matches. No human rater needed for millions of math problems. Models trained this way became dramatically better at reasoning.
Model generates code. Unit tests run automatically. Pass = reward. Fail = penalty. The model learns to write code that actually works — not just code that looks plausible.
Why RLVR is exciting right now: Scalable. Cheap. No human bottleneck. The "reasoning model" revolution (o1, o3, R1, QwQ) is largely built on RLVR applied to math and logic problems.
Two tools for the same job — choose based on whether correctness is measurable
| Property | RLHF Reinforcement Learning from Human Feedback |
RLVR Reinforcement Learning from Verifiable Rewards |
|---|---|---|
| Reward source | Human raters compare outputs | Automated checker (test runner, verifier) |
| Scales to… | How many humans you can hire | How much compute you have |
| Best for | Open-ended quality: helpfulness, tone, safety, creativity | Objectively checkable tasks: math, code, logic, games |
| Key risk | Reward model can be gamed; human biases get learned | Reward hacking — passes tests but misses intent |
| Examples | ChatGPT, Claude, Gemini alignment | o1/o3, DeepSeek-R1, AlphaCode |
| Often used | Together — RLHF first, then RLVR for specific capabilities | Together — RLHF first, then RLVR for specific capabilities |
RL agents find the shortest path to a high score — not necessarily the intended one
The agent discovers it can score high without doing what you actually wanted.
"When a measure becomes a target, it ceases to be a good measure."
The moment you optimize directly for a proxy metric, the proxy stops accurately measuring what you cared about.
This is one of the core problems in AI alignment — and why simply "optimizing harder" doesn't always make AI better.
Why this matters for AI literacy: When you use an AI tool, ask yourself — what reward function was it optimized for? Engagement? User satisfaction scores? Test performance? That answer shapes everything the model does.
Modern LLMs aren't just trained on text — they're aligned using RL
Now that we understand how AI learns, let's understand how it learns language
Natural Language Processing — teaching computers to understand and generate human language. Tokenization, embeddings, attention.
The architecture that changed everything. "Attention Is All You Need" (2017). Why transformers replaced older sequence models.
How GPT, Claude, and Gemini work. Pre-training, fine-tuning, prompting, capabilities and limits. Connecting back to RL alignment.
Everything you've learned about RL and alignment lives inside every LLM you use. The next module is the payoff — how it all comes together.