MODULE 15
Neural Networks and Deep Learning
Backpropagation derived and computed by hand, then optimisers, CNNs, RNNs, autoencoders, VAEs, GANs and diffusion models.
38 lessons~18h reading
- 0126 min
From Perceptron to Deep Networks
BeginnerComing soonWhy stacking linear layers is pointless, what nonlinearity buys, and the anatomy of a deep network.
Assumes: Multi-Layer Perceptrons and Feed-Forward Networks
- 0224 min
The Universal Approximation Theorem
AdvancedComing soonWhat the theorem actually promises, what it does not, and why depth beats width in practice.
Assumes: From Perceptron to Deep Networks
- 0328 min
Forward Propagation
BeginnerComing soonLayer-by-layer computation with explicit matrix shapes, batching, and a full numeric forward pass.
Assumes: From Perceptron to Deep Networks
- 0436 min
Backpropagation
AdvancedComing soonThe algorithm derived from the multivariable chain rule, as a computation graph and in matrix form.
Assumes: Forward Propagation · The Multivariable Chain Rule
- 0534 min
Backpropagation: A Complete Numeric Example
AdvancedComing soonEvery number in a 2-2-1 network for one full training step, forward and backward, verified by finite differences.
Assumes: Backpropagation
- 0630 min
Activation Functions
BeginnerComing soonSigmoid, tanh, ReLU, Leaky ReLU, ELU, GELU, SiLU and softmax, with derivatives and saturation behaviour.
Assumes: Forward Propagation
- 0730 min
Loss Functions
IntermediateComing soonMSE, MAE, Huber, binary and categorical cross-entropy, KL divergence, and matching loss to task.
Assumes: Activation Functions
- 0826 min
Weight Initialisation
AdvancedComing soonWhy zeros and large randoms both fail; Xavier/Glorot and He initialisation derived from variance analysis.
Assumes: Backpropagation
- 0928 min
SGD and Mini-Batching
IntermediateComing soonBatch, stochastic and mini-batch gradient descent, batch-size effects, and the noise-as-regulariser view.
Assumes: Gradient Descent
- 1024 min
Momentum and Nesterov Acceleration
IntermediateComing soonAccumulating velocity through ravines, and the look-ahead correction of Nesterov momentum.
Assumes: SGD and Mini-Batching
- 1132 min
Adaptive Optimisers: AdaGrad to AdamW
AdvancedComing soonPer-parameter learning rates, AdaGrad's decay problem, RMSProp, Adam's bias correction, and AdamW's decoupled decay.
Assumes: Momentum and Nesterov Acceleration
- 1224 min
Learning Rate Schedules
IntermediateComing soonStep, exponential and cosine decay, warmup, cyclical rates, and the LR-range test.
Assumes: Adaptive Optimisers: AdaGrad to AdamW
- 1326 min
Vanishing and Exploding Gradients
AdvancedComing soonWhy gradients decay or blow up with depth, diagnosis by norm tracking, and clipping.
Assumes: Weight Initialisation
- 1430 min
Batch Normalisation
AdvancedComing soonNormalising activations, learnable scale and shift, train/inference discrepancy, and the running statistics trap.
Assumes: Vanishing and Exploding Gradients
- 1524 min
Layer, Group and RMS Normalisation
AdvancedComing soonNormalising across features instead of the batch, and why transformers use LayerNorm and RMSNorm.
Assumes: Batch Normalisation
- 1624 min
Dropout
IntermediateComing soonRandom unit deactivation, inverted dropout at inference, and the ensemble interpretation.
Assumes: Regularisation
- 1724 min
Early Stopping and Data Augmentation
BeginnerComing soonPatience and restoration, plus augmentation strategies for images, text and tabular data.
Assumes: Dropout
- 1832 min
The Convolution Operation
IntermediateComing soonKernels, stride, padding and dilation; output-size arithmetic and a convolution computed by hand.
Assumes: Forward Propagation
- 1924 min
Pooling and Receptive Fields
IntermediateComing soonMax and average pooling, global pooling, and computing the receptive field of a deep stack.
Assumes: The Convolution Operation
- 2030 min
CNN Architectures
IntermediateComing soonLeNet, AlexNet, VGG, Inception and EfficientNet, and the design principles each introduced.
Assumes: Pooling and Receptive Fields
- 2128 min
ResNets and Skip Connections
AdvancedComing soonThe degradation problem, residual blocks, and why identity shortcuts make gradients flow.
Assumes: CNN Architectures
- 2228 min
Transfer Learning and Fine-Tuning CNNs
IntermediateComing soonFeature extraction vs fine-tuning, layer freezing, discriminative learning rates and domain shift.
Assumes: ResNets and Skip Connections
- 2332 min
Object Detection
AdvancedComing soonIoU, anchor boxes, non-maximum suppression, R-CNN family, YOLO, SSD and mAP evaluation.
Assumes: Transfer Learning and Fine-Tuning CNNs
- 2430 min
Semantic and Instance Segmentation
AdvancedComing soonFully convolutional networks, U-Net, transposed convolution, Mask R-CNN and Dice loss.
Assumes: Object Detection
- 2530 min
Recurrent Neural Networks
IntermediateComing soonHidden state recurrence, weight sharing across time, and the sequence-modelling task types.
Assumes: Backpropagation
- 2630 min
Backpropagation Through Time
AdvancedComing soonUnrolling the graph, gradient accumulation across timesteps, truncated BPTT, and why long dependencies fail.
Assumes: Recurrent Neural Networks
- 2734 min
Long Short-Term Memory
AdvancedComing soonCell state, forget/input/output gates, the constant error carousel, and a gate-by-gate numeric trace.
Assumes: Backpropagation Through Time
- 2822 min
Gated Recurrent Units
AdvancedComing soonUpdate and reset gates, the parameter saving over LSTM, and empirical comparisons.
Assumes: Long Short-Term Memory
- 2930 min
Sequence-to-Sequence Models
AdvancedComing soonEncoder–decoder architecture, the fixed-vector bottleneck, teacher forcing and beam search.
Assumes: Gated Recurrent Units
- 3028 min
Autoencoders
IntermediateComing soonEncoder, bottleneck and decoder; reconstruction loss, and the relationship to PCA.
Assumes: Backpropagation · Principal Component Analysis
- 3126 min
Denoising, Sparse and Contractive Autoencoders
AdvancedComing soonCorruption-based training, sparsity penalties, contractive regularisation and representation quality.
Assumes: Autoencoders
- 3236 min
Variational Autoencoders
AdvancedComing soonLatent variable modelling, the ELBO derived in full, the reparameterisation trick and posterior collapse.
Assumes: Denoising, Sparse and Contractive Autoencoders · Expectation–Maximisation
- 3332 min
Generative Adversarial Networks
AdvancedComing soonThe minimax game, discriminator and generator objectives, and the optimal-discriminator analysis.
Assumes: Variational Autoencoders · Zero-Sum Games and the Minimax Theorem
- 3430 min
GAN Training Dynamics and Variants
AdvancedComing soonMode collapse, vanishing discriminator gradients, DCGAN, WGAN with gradient penalty, StyleGAN and FID.
Assumes: Generative Adversarial Networks
- 3536 min
Diffusion Models
AdvancedComing soonForward noising and reverse denoising, the training objective, DDPM vs DDIM, and classifier-free guidance.
Assumes: Variational Autoencoders
- 3630 min
Self-Supervised and Contrastive Learning
AdvancedComing soonPretext tasks, InfoNCE, SimCLR, MoCo, BYOL and CLIP, and why negatives matter.
Assumes: Denoising, Sparse and Contractive Autoencoders
- 3730 min
Debugging Deep Networks
IntermediateComing soonA systematic checklist: overfit one batch, check shapes and gradients, and read loss curves diagnostically.
Assumes: Learning Rate Schedules
- 3834 min
PyTorch and TensorFlow Side by Side
IntermediateComing soonThe same network implemented in both frameworks: tensors, autograd, modules, training loops and deployment.
Assumes: Debugging Deep Networks