Skip to content
VibeFormer

MODULE 13

Supervised Learning

Every supervised algorithm on the syllabus, each derived from its objective, traced on small numeric data, then coded from scratch.

25 lessons~13h reading

  1. 01

    Simple Linear Regression

    BeginnerComing soon

    Fitting a line by least squares, deriving the closed-form slope and intercept, and interpreting them.

    Assumes: Least Squares and the Normal Equations

    30 min
  2. 02

    Multiple Linear Regression

    IntermediateComing soon

    The matrix formulation, the normal equations, multicollinearity, and interpreting partial coefficients.

    Assumes: Simple Linear Regression

    32 min
  3. 03

    Regression Assumptions and Diagnostics

    IntermediateComing soon

    Linearity, independence, homoscedasticity and normality; residual plots, and what to do when each fails.

    Assumes: Multiple Linear Regression · Regression Inference and Diagnostics

    30 min
  4. 04

    Polynomial and Basis Expansion Regression

    IntermediateComing soon

    Modelling curvature with linear machinery, basis functions, splines, and the overfitting cliff.

    Assumes: Multiple Linear Regression

    26 min
  5. 05

    Ridge Regression

    IntermediateComing soon

    The L2-penalised objective, its closed-form solution, the effect on the SVD spectrum, and choosing lambda.

    Assumes: Regularisation · Multiple Linear Regression

    30 min
  6. 06

    Lasso and Elastic Net

    AdvancedComing soon

    L1 geometry and sparsity, coordinate descent, the soft-thresholding operator, and elastic net as a hybrid.

    Assumes: Ridge Regression · Proximal Gradient Methods

    30 min
  7. 07

    Logistic Regression

    IntermediateComing soon

    The sigmoid, log-odds, the cross-entropy objective derived from MLE, and gradient-based fitting.

    Assumes: Maximum Likelihood Estimation · Gradient Descent

    34 min
  8. 08

    Multinomial Logistic Regression

    AdvancedComing soon

    Softmax, one-vs-rest versus multinomial formulations, and the gradient of softmax cross-entropy.

    Assumes: Logistic Regression

    28 min
  9. 09

    Linear Discriminant Analysis

    AdvancedComing soon

    Fisher's criterion, the generative Gaussian view, the shared-covariance assumption, and LDA for dimensionality reduction.

    Assumes: The Multivariate Normal Distribution · Eigenvalues and Eigenvectors

    32 min
  10. 10

    Quadratic Discriminant Analysis

    AdvancedComing soon

    Relaxing the equal-covariance assumption, the quadratic boundary, and the parameter-count trade-off.

    Assumes: Linear Discriminant Analysis

    22 min
  11. 11

    Naive Bayes Classifiers

    BeginnerComing soon

    The conditional independence assumption, Gaussian/multinomial/Bernoulli variants, and Laplace smoothing.

    Assumes: Bayes' Theorem

    32 min
  12. 12

    k-Nearest Neighbours

    BeginnerComing soon

    Instance-based learning, distance metrics, choosing k, weighting, and the cost of prediction.

    Assumes: The Curse of Dimensionality

    28 min
  13. 13

    Decision Trees: Entropy and Gini

    BeginnerComing soon

    Recursive partitioning, information gain, gain ratio and Gini impurity, with a tree built by hand.

    34 min
  14. 14

    CART, Pruning and Regression Trees

    IntermediateComing soon

    Binary splits, cost-complexity pruning, regression trees, and surrogate splits for missing values.

    Assumes: Decision Trees: Entropy and Gini

    28 min
  15. 15

    The Perceptron

    BeginnerComing soon

    The update rule, the convergence theorem for separable data, and the XOR limitation.

    26 min
  16. 16

    Support Vector Machines

    AdvancedComing soon

    Maximum-margin classification, the geometry of the margin, hard and soft margins, and the hinge loss.

    Assumes: The Perceptron · Convex Sets and Convex Functions

    34 min
  17. 17

    The SVM Dual Problem

    AdvancedComing soon

    Deriving the dual via Lagrangian duality, KKT conditions, and why only support vectors have nonzero multipliers.

    Assumes: Support Vector Machines · KKT Conditions · Linear Programming Duality

    32 min
  18. 18

    Kernel Methods

    AdvancedComing soon

    The kernel trick, Mercer's condition, polynomial and RBF kernels, and implicit feature spaces.

    Assumes: The SVM Dual Problem

    32 min
  19. 19

    Multi-Layer Perceptrons and Feed-Forward Networks

    IntermediateComing soon

    Stacking layers to defeat XOR, the forward pass, and why nonlinearity is essential — the bridge to deep learning.

    Assumes: The Perceptron

    30 min
  20. 20

    Bagging and Random Forests

    IntermediateComing soon

    Bootstrap aggregation, variance reduction arithmetic, feature subsampling, and out-of-bag error.

    Assumes: CART, Pruning and Regression Trees · Resampling: Bootstrap and Permutation

    32 min
  21. 21

    AdaBoost

    AdvancedComing soon

    Sequential reweighting, the derivation of the alpha coefficients, and the exponential-loss interpretation.

    Assumes: Bagging and Random Forests

    32 min
  22. 22

    Gradient Boosting

    AdvancedComing soon

    Boosting as gradient descent in function space, residual fitting, shrinkage and subsampling.

    Assumes: AdaBoost

    34 min
  23. 23

    XGBoost, LightGBM and CatBoost

    AdvancedComing soon

    Second-order objectives, regularised tree growth, histogram binning, leaf-wise growth and tuning strategy.

    Assumes: Gradient Boosting

    32 min
  24. 24

    Stacking and Blending

    AdvancedComing soon

    Meta-learners over base-model predictions, out-of-fold construction, and leakage avoidance.

    Assumes: Gradient Boosting

    24 min
  25. 25

    Model Interpretability

    AdvancedComing soon

    Coefficients, permutation importance, partial dependence, LIME and SHAP, with their assumptions and failure modes.

    Assumes: XGBoost, LightGBM and CatBoost · The Shapley Value

    32 min