MODULE 13
Supervised Learning
Every supervised algorithm on the syllabus, each derived from its objective, traced on small numeric data, then coded from scratch.
25 lessons~13h reading
- 0130 min
Simple Linear Regression
BeginnerComing soonFitting a line by least squares, deriving the closed-form slope and intercept, and interpreting them.
Assumes: Least Squares and the Normal Equations
- 0232 min
Multiple Linear Regression
IntermediateComing soonThe matrix formulation, the normal equations, multicollinearity, and interpreting partial coefficients.
Assumes: Simple Linear Regression
- 0330 min
Regression Assumptions and Diagnostics
IntermediateComing soonLinearity, independence, homoscedasticity and normality; residual plots, and what to do when each fails.
Assumes: Multiple Linear Regression · Regression Inference and Diagnostics
- 0426 min
Polynomial and Basis Expansion Regression
IntermediateComing soonModelling curvature with linear machinery, basis functions, splines, and the overfitting cliff.
Assumes: Multiple Linear Regression
- 0530 min
Ridge Regression
IntermediateComing soonThe L2-penalised objective, its closed-form solution, the effect on the SVD spectrum, and choosing lambda.
Assumes: Regularisation · Multiple Linear Regression
- 0630 min
Lasso and Elastic Net
AdvancedComing soonL1 geometry and sparsity, coordinate descent, the soft-thresholding operator, and elastic net as a hybrid.
Assumes: Ridge Regression · Proximal Gradient Methods
- 0734 min
Logistic Regression
IntermediateComing soonThe sigmoid, log-odds, the cross-entropy objective derived from MLE, and gradient-based fitting.
Assumes: Maximum Likelihood Estimation · Gradient Descent
- 0828 min
Multinomial Logistic Regression
AdvancedComing soonSoftmax, one-vs-rest versus multinomial formulations, and the gradient of softmax cross-entropy.
Assumes: Logistic Regression
- 0932 min
Linear Discriminant Analysis
AdvancedComing soonFisher's criterion, the generative Gaussian view, the shared-covariance assumption, and LDA for dimensionality reduction.
Assumes: The Multivariate Normal Distribution · Eigenvalues and Eigenvectors
- 1022 min
Quadratic Discriminant Analysis
AdvancedComing soonRelaxing the equal-covariance assumption, the quadratic boundary, and the parameter-count trade-off.
Assumes: Linear Discriminant Analysis
- 1132 min
Naive Bayes Classifiers
BeginnerComing soonThe conditional independence assumption, Gaussian/multinomial/Bernoulli variants, and Laplace smoothing.
Assumes: Bayes' Theorem
- 1228 min
k-Nearest Neighbours
BeginnerComing soonInstance-based learning, distance metrics, choosing k, weighting, and the cost of prediction.
Assumes: The Curse of Dimensionality
- 1334 min
Decision Trees: Entropy and Gini
BeginnerComing soonRecursive partitioning, information gain, gain ratio and Gini impurity, with a tree built by hand.
- 1428 min
CART, Pruning and Regression Trees
IntermediateComing soonBinary splits, cost-complexity pruning, regression trees, and surrogate splits for missing values.
Assumes: Decision Trees: Entropy and Gini
- 1526 min
The Perceptron
BeginnerComing soonThe update rule, the convergence theorem for separable data, and the XOR limitation.
- 1634 min
Support Vector Machines
AdvancedComing soonMaximum-margin classification, the geometry of the margin, hard and soft margins, and the hinge loss.
Assumes: The Perceptron · Convex Sets and Convex Functions
- 1732 min
The SVM Dual Problem
AdvancedComing soonDeriving the dual via Lagrangian duality, KKT conditions, and why only support vectors have nonzero multipliers.
Assumes: Support Vector Machines · KKT Conditions · Linear Programming Duality
- 1832 min
Kernel Methods
AdvancedComing soonThe kernel trick, Mercer's condition, polynomial and RBF kernels, and implicit feature spaces.
Assumes: The SVM Dual Problem
- 1930 min
Multi-Layer Perceptrons and Feed-Forward Networks
IntermediateComing soonStacking layers to defeat XOR, the forward pass, and why nonlinearity is essential — the bridge to deep learning.
Assumes: The Perceptron
- 2032 min
Bagging and Random Forests
IntermediateComing soonBootstrap aggregation, variance reduction arithmetic, feature subsampling, and out-of-bag error.
Assumes: CART, Pruning and Regression Trees · Resampling: Bootstrap and Permutation
- 2132 min
AdaBoost
AdvancedComing soonSequential reweighting, the derivation of the alpha coefficients, and the exponential-loss interpretation.
Assumes: Bagging and Random Forests
- 2234 min
Gradient Boosting
AdvancedComing soonBoosting as gradient descent in function space, residual fitting, shrinkage and subsampling.
Assumes: AdaBoost
- 2332 min
XGBoost, LightGBM and CatBoost
AdvancedComing soonSecond-order objectives, regularised tree growth, histogram binning, leaf-wise growth and tuning strategy.
Assumes: Gradient Boosting
- 2424 min
Stacking and Blending
AdvancedComing soonMeta-learners over base-model predictions, out-of-fold construction, and leakage avoidance.
Assumes: Gradient Boosting
- 2532 min
Model Interpretability
AdvancedComing soonCoefficients, permutation importance, partial dependence, LIME and SHAP, with their assumptions and failure modes.
Assumes: XGBoost, LightGBM and CatBoost · The Shapley Value