CAI1000 · Supplemental Module

From Learning Rules
to Language Models

ML Types Review · Reinforcement Learning · RLHF & RLVR · Intro to NLP

CAI1000 — Introduction to Artificial Intelligence
Review

The Three Flavors of Machine Learning

Every ML system learns differently based on what kind of feedback it receives

Supervised

Learns from labeled examples. You provide input + correct answer. The model learns the mapping.

Email → spam/not spam
Image → cat/dog
Loan application → approve/deny

Unsupervised

Finds hidden patterns in data with no labels. The model groups or compresses on its own.

Customer segments
Topic clustering
Anomaly detection

Reinforcement

Learns by trial and error in an environment. Gets rewards for good actions, penalties for bad ones.

Game playing (AlphaGo)
Robot locomotion
Ad bidding systems

Most real-world AI systems combine multiple approaches — these are categories, not strict boxes.

Semi-Supervised Learning

When Labels Are Expensive

Most real-world data is unlabeled — labeling it costs time, money, and human expertise

The Labeling Problem

  • A hospital has 1M X-rays — but only 500 have been reviewed by a radiologist
  • Training purely supervised: only 500 examples → poor generalization
  • Ignoring the 999,500 unlabeled ones wastes most of the data

Semi-Supervised Solution

  • Train on the 500 labeled examples to get an initial model
  • Use that model to pseudo-label high-confidence unlabeled cases
  • Retrain on labeled + pseudo-labeled data — repeat
  • Dramatically improves with way less human effort
1%
Labeled data
Expert-annotated, expensive to produce
99%
Unlabeled data
Abundant, cheap, but no answers provided
Modern LLMs are pre-trained on massive unlabeled corpora (self-supervised) — then fine-tuned with labeled data. That's the connection to your next module!
Reinforcement Learning

Learning by Doing

An agent takes actions in an environment and learns from the reward it receives

Agent
The AI making decisions
① Action
→
←
② State + Reward
Environment
The world the agent lives in

State

What the agent can "see" right now — the current situation.
Chess: current board positions

Action

What the agent decides to do from the available options.
Chess: which piece to move where

Reward

The score — positive or negative — the environment gives back.
Chess: +1 for winning, -1 for losing

Reinforcement Learning

What Makes a Good Reward?

The reward signal is everything — the agent will optimize for exactly what you measure

Good Rewards Are…

  • Clear and unambiguous
  • Directly tied to the real goal
  • Available frequently enough to learn from
  • Resistant to "gaming"

Bad Rewards Lead To…

  • Reward hacking — finding shortcuts that score high but miss the point
  • Unintended behavior that's hard to predict
  • Systems that are technically correct but socially wrong

Classic Examples

Boat racing game Drove in circles collecting power-ups instead of finishing
Robot arm Rewarded for fast hand speed — learned to vibrate violently
Social media feed Optimized for clicks → outrage became the strategy
AlphaGo Rewarded for winning → discovered strategies humans never imagined

"You get what you measure" — the alignment problem starts here.

Reinforcement Learning

Explore or Exploit?

Every RL agent faces this fundamental tradeoff on every step

Exploitation

Use what you already know to get the best reward you can right now.

  • Stick with strategies that worked before
  • Safe and predictable
  • Risk: you might be stuck in a local optimum

You always go to the same restaurant because it's good. But you'll never find a better one.

Exploration

Try new things to gather information, even at the cost of short-term rewards.

  • Discover strategies not yet tried
  • Necessary for long-term improvement
  • Risk: short-term performance suffers

You try a new restaurant each week. Sometimes terrible — but once in a while, life-changing.

RL algorithms use strategies like ε-greedy (explore randomly X% of the time) or decaying exploration (explore a lot early, exploit later as confidence grows) to balance both sides.

ML Comparison

RL vs. Supervised Learning

Two very different kinds of feedback loops

Labeled Dataset Correct answers provided
→
Model Training
→
Fixed Prediction

Supervised: train once, deploy, done. No feedback from the real world after training.

Agent Acts
→
Environment
Responds
→
Reward Signal
→
Policy Updates
↩

RL: continuous loop — the agent keeps learning from consequences, not from pre-labeled answers.

Feedback typeSupervised: correct answer
When does it learn?Before deployment
Best forFixed mappings
Feedback typeRL: reward/penalty
When does it learn?During interaction
Best forSequential decisions
Use Cases

Where Reinforcement Learning Shines

Any domain where you make sequences of decisions and can define a reward

Games & Strategy

AlphaGo, AlphaStar (StarCraft II), OpenAI Five (Dota 2). Reward: winning. Discovered strategies no human had ever played.

Robotics

Teaching robots to walk, grasp, and navigate. Reward: staying upright, reaching targets, picking up objects without dropping.

Recommendations

Netflix, YouTube, Spotify. Reward: watch time, clicks, satisfaction ratings. Each user interaction is feedback for the next suggestion.

Drug Discovery

Designing new molecule structures. Reward: predicted binding affinity, low toxicity. Explores chemical space faster than human chemists.

Common thread: sequential decisions + measurable outcomes + lots of trial opportunities. If any of these is missing, RL gets much harder.

RLHF

Reinforcement Learning from Human Feedback

When you can't write down the reward — you use humans to teach what "good" means

Step 1
Supervised Fine-Tuning Train on high-quality human-written examples. Gives the model a foundation.
→
Step 2
Human Preference Data Show humans pairs of AI responses. Ask: "Which is better?" Collect thousands of comparisons.
→
Step 3
Train Reward Model A second AI learns to predict which responses humans would prefer. This becomes the reward signal.
→
Step 4
RL Fine-Tuning (PPO) Optimize the main model to maximize the reward model's score. Aligns output with human values.

Why RLHF matters for LLMs: You can't write a formula for "helpful, harmless, honest." RLHF lets humans teach those properties directly. GPT-4, Claude, and Gemini all use versions of this process.

RLHF

Before and After RLHF

The same underlying model behaves very differently once human preferences are trained in

Without RLHF (base model)
How do I stay motivated to exercise?
Exercise is a physical activity that involves moving the body. Common forms include running, swimming, and lifting weights. The human body benefits from regular exercise through improved cardiovascular health and...
...weight management strategies vary by individual and may include caloric restriction alongside physical activity...

Technically correct — but robotic, generic, and doesn't actually answer the question.

With RLHF (aligned model)
How do I stay motivated to exercise?
A few things that really help: start ridiculously small — 5 minutes counts. Tie it to something you already do (walk after lunch). And don't rely on motivation — build a routine so it happens even when you don't feel like it. What's getting in the way for you?

Helpful, practical, conversational, and asks a follow-up. This is what human raters preferred.

The model learned this style from thousands of pairwise human comparisons — not explicit instructions.

RLVR

Reinforcement Learning from Verifiable Rewards

What if you could check the answer automatically — no human needed?

The Key Insight

For some tasks, correctness is objective and checkable:

  • Did the code run without errors and pass the test cases?
  • Is the math answer numerically correct?
  • Does the logical proof follow valid inference rules?
  • Did the AI win or lose the game?

These are "verifiable rewards" — no human judgment needed, just a checker.

RLVR for Math (DeepSeek-R1, o1)

Model generates a chain-of-thought answer. A verifier checks if the final number matches. No human rater needed for millions of math problems. Models trained this way became dramatically better at reasoning.

RLVR for Code (AlphaCode, GitHub Copilot)

Model generates code. Unit tests run automatically. Pass = reward. Fail = penalty. The model learns to write code that actually works — not just code that looks plausible.

Why RLVR is exciting right now: Scalable. Cheap. No human bottleneck. The "reasoning model" revolution (o1, o3, R1, QwQ) is largely built on RLVR applied to math and logic problems.

Comparison

RLHF vs. RLVR

Two tools for the same job — choose based on whether correctness is measurable

Property RLHF
Reinforcement Learning from Human Feedback
RLVR
Reinforcement Learning from Verifiable Rewards
Reward source Human raters compare outputs Automated checker (test runner, verifier)
Scales to… How many humans you can hire How much compute you have
Best for Open-ended quality: helpfulness, tone, safety, creativity Objectively checkable tasks: math, code, logic, games
Key risk Reward model can be gamed; human biases get learned Reward hacking — passes tests but misses intent
Examples ChatGPT, Claude, Gemini alignment o1/o3, DeepSeek-R1, AlphaCode
Often used Together — RLHF first, then RLVR for specific capabilities Together — RLHF first, then RLVR for specific capabilities
Limits & Alignment

When Optimization Goes Wrong

RL agents find the shortest path to a high score — not necessarily the intended one

Reward Hacking

The agent discovers it can score high without doing what you actually wanted.

  • RLHF model learns to sound confident and pleasant even when wrong
  • Code model writes code that passes tests by hardcoding test answers
  • Recommendation system maximizes watch time by promoting outrage

Goodhart's Law

"When a measure becomes a target, it ceases to be a good measure."

The moment you optimize directly for a proxy metric, the proxy stops accurately measuring what you cared about.

This is one of the core problems in AI alignment — and why simply "optimizing harder" doesn't always make AI better.

Why this matters for AI literacy: When you use an AI tool, ask yourself — what reward function was it optimized for? Engagement? User satisfaction scores? Test performance? That answer shapes everything the model does.

Connection

RL Powers the AI You Use Every Day

Modern LLMs aren't just trained on text — they're aligned using RL

Pre-training Predict next token on trillions of words. Self-supervised. No labels needed.
→
SFT Supervised fine-tuning on high-quality Q&A and instruction-following examples.
→
RLHF Human raters compare outputs. Reward model trained. PPO optimization for helpfulness + safety.
→
RLVR For reasoning tasks: math, code, logic. Verifiable reward signals boost capability.
→
ChatGPT / Claude / Gemini Helpful, harmless, capable assistants
Step 1
Learn language structure from massive text (no RL yet)
Step 2
Learn human values through RLHF — the alignment step
Step 3
Boost specific skills with RLVR — the reasoning step
Coming Up Next

NLP & Language Models

Now that we understand how AI learns, let's understand how it learns language

What is NLP?

Natural Language Processing — teaching computers to understand and generate human language. Tokenization, embeddings, attention.

Transformers

The architecture that changed everything. "Attention Is All You Need" (2017). Why transformers replaced older sequence models.

Large Language Models

How GPT, Claude, and Gemini work. Pre-training, fine-tuning, prompting, capabilities and limits. Connecting back to RL alignment.

Everything you've learned about RL and alignment lives inside every LLM you use. The next module is the payoff — how it all comes together.

1 / 15
CAI1000 — RL & NLP Bridge
← → or click to navigate