warming up your workspace

Deep Learning

Build an understanding of neural networks through Python. Follow regression, training, convolutions, and attention towards a small GPT model.

What helps

Start with Python Foundations if needed. Algebra, vectors, derivatives, and probability help explain how the later models learn.

Python pathway

Programming Foundations / Practice rooms / Track curriculum and enrollment

  1. Foundations: From Regression to the Neuron

    Start with a model that predicts numbers and gradient descent that adjusts its parameters. Fit linear regression, work with multiple features, and build weighted sums and activation functions. A sigmoid neuron with a binary classification loss gives logistic regression. Finish by training and evaluating a neuron on the AND truth table.

    • Linear Regression: 5 lessons
    • Gradient Descent in Practice: 5 lessons
    • The Artificial Neuron: 5 lessons
    • Logistic Regression: 5 lessons
    • Capstone: A Neuron Learns a Logic Gate: 5 lessons
  2. Neural Networks

    Combine neurons into dense layers and layers into a multilayer perceptron. Write matrix-based forward passes, use softmax for class probabilities, and construct a hidden-layer network for XOR, which a single linear decision boundary cannot separate. Finish with reusable initialization, prediction and accuracy reporting; training these networks comes next.

    • The Dense Layer: 5 lessons
    • The Multilayer Perceptron: 5 lessons
    • Multi-class Output: 5 lessons
    • The Power of Depth: 5 lessons
    • Capstone: A Reusable Network: 5 lessons
  3. Backpropagation and Autograd

    Backpropagation computes loss gradients by applying the chain rule backward through a computation graph; an optimizer then uses those gradients to update parameters. Differentiate a neuron and a two-layer network, compare derivatives with finite differences, build a scalar autograd engine, and train an XOR network while retaining its weights and loss history.

    • The Chain Rule: 5 lessons
    • Backprop Through a Neuron: 5 lessons
    • Backprop Through a Network: 5 lessons
    • A Tiny Autograd Engine: 5 lessons
    • Capstone: Learn XOR with Backprop: 5 lessons
  4. Training Neural Networks

    Build losses that score predictions and optimizers that update parameters: SGD, momentum, RMSprop and Adam. Combine forward and backward passes into mini-batch training with persistent optimizer state and epoch histories. Train a small regression network on synthetic data and measure its progress; a chosen update rule does not guarantee improvement on every step.

    • Loss Functions: 5 lessons
    • Gradient Descent and SGD: 5 lessons
    • Momentum: 5 lessons
    • Adaptive Optimizers: 5 lessons
    • The Training Loop: 5 lessons
  5. Deep Classification

    Train a multi-class network on synthetic clusters and distinguish training accuracy from performance on held-out data. Implement stable softmax cross-entropy, L2 regularization, dropout, validation splits, early stopping and confusion-based metrics. Assemble a classifier that retains the best validation parameters and reports measured results; regularization benefits must be evaluated.

    • The Softmax Classifier: 5 lessons
    • Training a Classifier: 5 lessons
    • Regularization: 5 lessons
    • Evaluating a Model: 5 lessons
    • Capstone: A Regularized Classifier: 5 lessons
  6. Convolutional Networks

    Convolutional networks build in local connectivity and shared filters for spatial inputs. Implement the unflipped sliding-kernel convention used in CNNs, inspect hand-designed filters, pool feature maps and differentiate a convolutional layer. Train a tiny CNN on synthetic vertical and horizontal bars, inspect its learned filters and evaluate held-out images.

    • Convolution: 5 lessons
    • Filters and Feature Maps: 5 lessons
    • Pooling: 5 lessons
    • The Convolutional Layer: 5 lessons
    • Capstone: Training a CNN: 5 lessons
  7. Sequences and Recurrent Networks

    Process ordered inputs one step at a time with a recurrent hidden state, a compressed representation rather than a perfect memory. Build the RNN cell and sequence forward pass, derive backpropagation through time, examine gradient decay and clipping, and train a character model. Finish with reusable training, reporting, greedy generation and temperature sampling.

    • Sequences and State: 5 lessons
    • The RNN Forward Pass: 5 lessons
    • Backpropagation Through Time: 5 lessons
    • Training an RNN: 5 lessons
    • Capstone: A Char-RNN that Generates: 5 lessons
  8. Tokenization and Embeddings

    Turn raw text into stable token IDs and learn dense vectors for those IDs. Build character, word and UTF-8 byte tokenizers, embedding lookup and accumulated gradients, cosine similarity and vector analogy calculations. Train a small bag-of-embeddings sentiment classifier and inspect its predictions and neighbors; useful semantic geometry depends on the data and objective, not on lookup alone.

    • Tokenization: 5 lessons
    • Embeddings: 5 lessons
    • Vector Similarity: 5 lessons
    • Learning Embeddings: 5 lessons
    • Capstone: A Text Classifier: 5 lessons
  9. Attention and Transformers

    Build scaled dot-product attention, causal masks, sinusoidal position information and multiple attention heads. Combine output projections, residual connections, layer normalization and feed-forward layers into stacked transformer blocks and a reusable next-token forward model. Full attention compares token pairs and has quadratic sequence cost at fixed width; generation still proceeds one token at a time.

    • Scaled Dot-Product Attention: 5 lessons
    • Self-Attention: 5 lessons
    • Multi-Head Attention and Position: 5 lessons
    • The Transformer Block: 5 lessons
    • Capstone: A Self-Attention Language Model: 5 lessons
  10. Capstone: A Tiny GPT

    Build and train an educational character-level autoregressive transformer. Compose embeddings and positions, strict causal attention, residual feed-forward layers and an output head; differentiate all eleven parameter arrays and check them numerically. Train on a tiny corpus with Adam, retain optimizer state and history, and provide reporting, greedy generation and temperature sampling. This single-block training model omits layer normalization and does not claim the architecture or capabilities of a production GPT.

    • The GPT Forward Pass: 5 lessons
    • Backpropagation: The Top Layers: 5 lessons
    • Backpropagation: Attention: 5 lessons
    • Training the GPT: 5 lessons
    • Generating Text: 5 lessons