Deep Learning
Build an understanding of neural networks through Python. Follow regression, training, convolutions, and attention towards a small GPT model.
What helps
Start with Python Foundations if needed. Algebra, vectors, derivatives, and probability help explain how the later models learn.
Python pathway
Programming Foundations / Practice rooms / Track curriculum and enrollment
Foundations: From Regression to the Neuron
Start with a model that predicts numbers and gradient descent that adjusts its parameters. Fit linear regression, work with multiple features, and build weighted sums and activation functions. A sigmoid neuron with a binary classification loss gives logistic regression. Finish by training and evaluating a neuron on the AND truth table.
- Linear Regression: 5 lessons
- Gradient Descent in Practice: 5 lessons
- The Artificial Neuron: 5 lessons
- Logistic Regression: 5 lessons
- Capstone: A Neuron Learns a Logic Gate: 5 lessons
Neural Networks
Combine neurons into dense layers and layers into a multilayer perceptron. Write matrix-based forward passes, use softmax for class probabilities, and construct a hidden-layer network for XOR, which a single linear decision boundary cannot separate. Finish with reusable initialization, prediction and accuracy reporting; training these networks comes next.
- The Dense Layer: 5 lessons
- The Multilayer Perceptron: 5 lessons
- Multi-class Output: 5 lessons
- The Power of Depth: 5 lessons
- Capstone: A Reusable Network: 5 lessons
Backpropagation and Autograd
Backpropagation computes loss gradients by applying the chain rule backward through a computation graph; an optimizer then uses those gradients to update parameters. Differentiate a neuron and a two-layer network, compare derivatives with finite differences, build a scalar autograd engine, and train an XOR network while retaining its weights and loss history.
- The Chain Rule: 5 lessons
- Backprop Through a Neuron: 5 lessons
- Backprop Through a Network: 5 lessons
- A Tiny Autograd Engine: 5 lessons
- Capstone: Learn XOR with Backprop: 5 lessons
Training Neural Networks
Build losses that score predictions and optimizers that update parameters: SGD, momentum, RMSprop and Adam. Combine forward and backward passes into mini-batch training with persistent optimizer state and epoch histories. Train a small regression network on synthetic data and measure its progress; a chosen update rule does not guarantee improvement on every step.
- Loss Functions: 5 lessons
- Gradient Descent and SGD: 5 lessons
- Momentum: 5 lessons
- Adaptive Optimizers: 5 lessons
- The Training Loop: 5 lessons
Deep Classification
Train a multi-class network on synthetic clusters and distinguish training accuracy from performance on held-out data. Implement stable softmax cross-entropy, L2 regularization, dropout, validation splits, early stopping and confusion-based metrics. Assemble a classifier that retains the best validation parameters and reports measured results; regularization benefits must be evaluated.
- The Softmax Classifier: 5 lessons
- Training a Classifier: 5 lessons
- Regularization: 5 lessons
- Evaluating a Model: 5 lessons
- Capstone: A Regularized Classifier: 5 lessons
Convolutional Networks
Convolutional networks build in local connectivity and shared filters for spatial inputs. Implement the unflipped sliding-kernel convention used in CNNs, inspect hand-designed filters, pool feature maps and differentiate a convolutional layer. Train a tiny CNN on synthetic vertical and horizontal bars, inspect its learned filters and evaluate held-out images.
- Convolution: 5 lessons
- Filters and Feature Maps: 5 lessons
- Pooling: 5 lessons
- The Convolutional Layer: 5 lessons
- Capstone: Training a CNN: 5 lessons
Sequences and Recurrent Networks
Process ordered inputs one step at a time with a recurrent hidden state, a compressed representation rather than a perfect memory. Build the RNN cell and sequence forward pass, derive backpropagation through time, examine gradient decay and clipping, and train a character model. Finish with reusable training, reporting, greedy generation and temperature sampling.
- Sequences and State: 5 lessons
- The RNN Forward Pass: 5 lessons
- Backpropagation Through Time: 5 lessons
- Training an RNN: 5 lessons
- Capstone: A Char-RNN that Generates: 5 lessons
Tokenization and Embeddings
Turn raw text into stable token IDs and learn dense vectors for those IDs. Build character, word and UTF-8 byte tokenizers, embedding lookup and accumulated gradients, cosine similarity and vector analogy calculations. Train a small bag-of-embeddings sentiment classifier and inspect its predictions and neighbors; useful semantic geometry depends on the data and objective, not on lookup alone.
- Tokenization: 5 lessons
- Embeddings: 5 lessons
- Vector Similarity: 5 lessons
- Learning Embeddings: 5 lessons
- Capstone: A Text Classifier: 5 lessons
Attention and Transformers
Build scaled dot-product attention, causal masks, sinusoidal position information and multiple attention heads. Combine output projections, residual connections, layer normalization and feed-forward layers into stacked transformer blocks and a reusable next-token forward model. Full attention compares token pairs and has quadratic sequence cost at fixed width; generation still proceeds one token at a time.
- Scaled Dot-Product Attention: 5 lessons
- Self-Attention: 5 lessons
- Multi-Head Attention and Position: 5 lessons
- The Transformer Block: 5 lessons
- Capstone: A Self-Attention Language Model: 5 lessons
Capstone: A Tiny GPT
Build and train an educational character-level autoregressive transformer. Compose embeddings and positions, strict causal attention, residual feed-forward layers and an output head; differentiate all eleven parameter arrays and check them numerically. Train on a tiny corpus with Adam, retain optimizer state and history, and provide reporting, greedy generation and temperature sampling. This single-block training model omits layer normalization and does not claim the architecture or capabilities of a production GPT.
- The GPT Forward Pass: 5 lessons
- Backpropagation: The Top Layers: 5 lessons
- Backpropagation: Attention: 5 lessons
- Training the GPT: 5 lessons
- Generating Text: 5 lessons