CAI1000 · Module 5 · NLP

How AI Learns
Language

One simple task, repeated trillions of times

Artificial intelligence is demonstrated by ?
Core Concept

Predict the next word.

That's the entire training objective. Everything else emerges from doing it at scale.

Wikipedia — Artificial Intelligence

Artificial intelligence is intelligence demonstrated by ? , as opposed to…

The model guesses

It predicts a probability distribution over every word in its vocabulary — which word is most likely to come next?

It gets corrected

The actual word is revealed. If the model was wrong, its weights are nudged slightly. Do this billions of times: you get ChatGPT.

Training Data

Any text becomes thousands of examples

One Wikipedia sentence → one sliding window → many (context, target) pairs

Wikipedia article — "Artificial Intelligence"

Artificial intelligence is intelligence demonstrated by machines , as opposed to the natural intelligence displayed by humans .

amber = context   purple = target to predict

Example 1 → Context: "Artificial" → Predict: "intelligence"
Example 2 → Context: "Artificial intelligence" → Predict: "is"
Example 3 → Context: "Artificial intelligence is" → Predict: "intelligence"
Example 4 → Context: "Artificial intelligence is intelligence" → Predict: "demonstrated"
Example 5 → Context: "Artificial intelligence is intelligence demonstrated" → Predict: "by"
Example 6 → Context: "Artificial intelligence is intelligence demonstrated by" → Predict: "machines" ← shown on slide 1
Model Output

Not a single guess — a ranked list

The model outputs a probability for every word in its vocabulary. Here are the top 5.

Context fed to model

Artificial intelligence is intelligence demonstrated by ?

The model has never been told what words mean — it has only seen which words tend to follow others, across billions of sentences.

GPT-2 top-5 predictions

1. machines
38.2% ✓ actual
2. humans
22.1%
3. computers
15.4%
4. robots
 9.8%
5. algorithms
 5.1%

The correct word "machines" gets the highest probability — the model has absorbed enough context to know this sentence is describing a technical concept, not a biological one.

Live Demo

Run it yourself — 30 lines of Python

GPT-2 (the predecessor to ChatGPT) is free, runs locally, and fits in ~500 MB

# install once: pip install transformers torch
$ python nlp_causal_demo.py
──────────────────────────────────────────────────────────────
Context: "Artificial intelligence is intelligence demonstrated by [?]"
Top 5 predictions:
  1. machines        38.2%  ████████████████████████████ ← actual word
  2. humans          22.1%  ████████████████
  3. computers       15.4%  ███████████
  4. robots           9.8%  ███████
  5. algorithms       5.1%  ███
──────────────────────────────────────────────────────────────
Context: "AI research has been defined as the field of study of [?]"
Top 5 predictions:
  1. intelligent     41.5%  ██████████████████████████████ ← actual word
  2. artificial      19.2%  ██████████████
  3. human           12.8%  █████████

Script file: nlp_causal_demo.py — open it in any Python environment or Google Colab

Training Pipeline

From raw text to a trained model

The same four steps, repeated for every token in the training data

Raw Text Wikipedia, books, web
→
Sliding Window context + target pairs
→
Model Guesses probability distribution
→
Compare to Actual measure the error
→
Adjust Weights backpropagation

The model never sees a grammar textbook, a dictionary, or rules about language.

It learns grammar, facts, reasoning, and style entirely by predicting what comes next, from patterns in text.

The Punchline

Scale turns word prediction into intelligence

The task is simple. The scale is incomprehensible. The emergent capability is surprising.

3T
tokens in GPT-3 training data
175B
parameters (weights) in GPT-3
1 task
predict the next token

What it had to learn to do the task

  • → Grammar & syntax (to predict plausible words)
  • → World knowledge (to predict factual words)
  • → Reasoning (to predict logically consistent words)
  • → Writing style (to predict stylistically appropriate words)

What emerged "for free"

  • → Code generation (code is text too)
  • → Translation (the model saw multilingual text)
  • → Summarization (it learned what words to skip)
  • → Question answering (it saw Q&A pairs in training)
1 / 7
CAI1000 · NLP · Causal Language Modeling
← → or click to navigate