One simple task, repeated trillions of times
That's the entire training objective. Everything else emerges from doing it at scale.
Wikipedia — Artificial Intelligence
It predicts a probability distribution over every word in its vocabulary — which word is most likely to come next?
The actual word is revealed. If the model was wrong, its weights are nudged slightly. Do this billions of times: you get ChatGPT.
One Wikipedia sentence → one sliding window → many (context, target) pairs
Wikipedia article — "Artificial Intelligence"
amber = context purple = target to predict
The model outputs a probability for every word in its vocabulary. Here are the top 5.
Context fed to model
The model has never been told what words mean — it has only seen which words tend to follow others, across billions of sentences.
GPT-2 top-5 predictions
The correct word "machines" gets the highest probability — the model has absorbed enough context to know this sentence is describing a technical concept, not a biological one.
GPT-2 (the predecessor to ChatGPT) is free, runs locally, and fits in ~500 MB
Script file: nlp_causal_demo.py — open it in any Python environment or Google Colab
The same four steps, repeated for every token in the training data
The model never sees a grammar textbook, a dictionary, or rules about language.
It learns grammar, facts, reasoning, and style entirely by predicting what comes next, from patterns in text.
The task is simple. The scale is incomprehensible. The emergent capability is surprising.