From Rules to Neurons

Understanding Deep Learning — and why it changed everything

CAI1000 · Module 3/4 Bridge · Supplemental
Course Context

You Already Know the Core Idea

Machine learning learns patterns from data — no explicit rules required.

Data examples & labels
→
ML Algorithm finds patterns
→
Model makes predictions

What you've seen so far

  • Supervised & unsupervised learning
  • Decision trees, regression, k-means
  • Train / test splits and evaluation
  • Features → predictions

What's coming in Modules 5 & 6

  • AI that understands language
  • AI that recognizes images
  • ChatGPT, DALL-E, self-driving cars
  • All powered by deep learning

This deck fills the gap between the two — what deep learning is, and why it's a different beast.

The Problem

Traditional ML Had a Hidden Bottleneck

A human expert had to decide what features to feed the model.

Traditional ML: feature engineering

  • 📋 Detect spam email?
  • → Expert manually picks: word counts, sender domain, punctuation rate, caps ratio…
  • 🖼️ Recognize a cat?
  • → Expert manually defines: ear shape, whisker angles, fur texture, eye position…
The problem: Writing these rules by hand is slow, expensive, and breaks on anything the expert didn't anticipate. What about a cat wearing a hat?
🧑‍💼

The Bottleneck

Human expertise required at every step.
Doesn't scale. Doesn't generalize.

The question researchers asked:

What if instead of telling the machine what to look for… we let it figure that out itself?

Core Concept

The Neural Network: Layers That Learn

Inspired by the brain — but a very simplified version of it.

Input Hidden Hidden Output

Input layer

Raw data goes in — pixel values, word tokens, numbers. No human hand-crafting required.

Hidden layers

The network automatically discovers which patterns matter. You don't program these — the training process figures them out.

Output layer

A prediction: "cat vs dog," "spam vs not spam," "translate this sentence."

How Layers Work

Each Layer Learns Something More Abstract

Imagine the network learning to recognize a human face from a photo.

Raw Pixels 255, 247, 201…
→
Edges & Lines layer 1–2
→
Shapes & Curves layer 3–5
→
Face Parts eyes, nose, mouth
→
"Person" final prediction
Nobody programmed "look for edges." The network learned that edges are useful because they kept showing up in faces during training. It discovered this pattern on its own.
"Deep" in deep learning means many layers — deeper networks can learn more abstract, complex features. A 200-layer network can recognize things a 3-layer network never could.

The same principle applies whether the input is images, text, audio, or sensor data — layers build progressively richer representations.

Training

How a Network Learns: Trial, Error, Adjust

No programmer writes the internal rules — the network finds them by making millions of mistakes.

Show example labeled training data
→
Make a guess forward pass
→
Measure error how wrong was it?
→
Adjust weights backpropagation
→
Repeat millions of times
Weights

The numbers on each connection. Training adjusts billions of these tiny dials.

Loss

A single number measuring how wrong the guess was. Training tries to make it as small as possible.

Epochs

One full pass through the training data. A model might train for hundreds of epochs over days or weeks.

The Breakthrough

Scale Is the Secret Ingredient

Neural networks existed for decades. What changed wasn't the idea — it was the scale.

2012
AlexNet wins ImageNet by a landslide — the "Big Bang" of modern deep learning
1,000×
More training data available in 2012 than in 1990, thanks to the internet
GPUs
Graphics cards originally built for video games turned out to be perfect for training neural nets in parallel

The "Scaling Law" insight

As you give a neural network more data and more computing power, it keeps getting better — reliably, predictably. This was not obvious, and it changed how everyone thinks about AI.

From AlexNet to GPT-4

AlexNet (2012) parameters 60 million
GPT-3 (2020) parameters 175 billion
GPT-4 (est.) parameters ~1.8 trillion

Each "parameter" is one weight in the network — one dial the training process tuned.

Comparison

ML vs. Deep Learning — Picking the Right Tool

Deep learning is not always better. It depends on what you have to work with.

Traditional ML — reach for this when…

  • 📊 You have a small or mid-size dataset
  • 🔍 You need to explain the decision to stakeholders
  • ⚡ You need fast training on a laptop
  • 📋 Your features are already well-understood
  • Decision Trees Regression K-Means

Deep Learning — reach for this when…

  • 🗂️ You have massive amounts of data
  • 🖼️ Your input is raw images, audio, or text
  • 💡 Nobody knows which features matter — let it learn
  • 🎯 You need state-of-the-art accuracy
  • CNNs Transformers LLMs
One more thing: Deep learning is a subset of machine learning — not a replacement. All deep learning is ML. Not all ML is deep learning. It's a spectrum, not two separate worlds.
Coming Up

This Is the Engine Under the Hood

The next two modules will make a lot more sense now that you know what's powering them.

Module 5 — Natural Language Processing

ChatGPT, Claude, Gemini — all powered by Transformer neural networks with billions of layers and hundreds of billions of weights.

  • → Pre-training: Learn from the entire internet (deep learning at scale)
  • → Fine-tuning: Specialize for conversations and instructions
  • → Each token you generate = one forward pass through the network

Module 6 — Computer Vision

Image recognition, object detection, medical imaging, self-driving cars — all powered by Convolutional Neural Networks (CNNs).

  • → Convolutional layers scan for local patterns (the abstraction ladder in action)
  • → Teachable Machine will let you train your own CNN
  • → That 2012 ImageNet moment? A CNN named AlexNet did it
Deep learning is not a different kind of AI — it's how modern AI gets its power. You now know what's happening inside the black box.
Summary

Where Deep Learning Fits in the AI Landscape

A nested view — every outer ring contains everything inside it.

Artificial Intelligence rule-based systems, expert systems, search, planning… Machine Learning decision trees, SVMs, k-means, regression… Deep Learning neural nets, CNNs, Transformers, LLMs ← You are here Modules 5 & 6

What to remember

Deep learning = neural networks with many layers that learn features automatically from large amounts of data.

Why it matters

It is the core technology behind virtually every headline-grabbing AI system of the past decade.

What's next

Modules 5 & 6 show you what deep learning looks like in action — for language and for images.

1 / 10
CAI1000 · Deep Learning Bridge
← → or click to navigate