MODULE 12
Machine Learning Foundations
The concepts every algorithm shares: risk minimisation, generalisation, the bias–variance trade-off, validation and metrics.
20 lessons~9h reading
- 0122 min
What Machine Learning Actually Is
BeginnerLearning from data versus explicit programming, the three paradigms, and an honest account of what ML cannot do.
- 0226 min
Formulating a Learning Problem
BeginnerInput and output spaces, hypothesis classes, loss functions, and turning a vague goal into an objective.
Assumes: What Machine Learning Actually Is
- 0328 min
Empirical Risk Minimisation
AdvancedTrue risk versus empirical risk, why we optimise a proxy, and the approximation–estimation decomposition.
Assumes: Formulating a Learning Problem
- 0428 min
Generalisation, Overfitting and Underfitting
BeginnerDiagnosing capacity problems from learning curves, with the classic polynomial-fit demonstration.
Assumes: Formulating a Learning Problem
- 0532 min
The Bias–Variance Trade-off
IntermediateFull algebraic decomposition of expected squared error into bias, variance and noise, with a simulation.
Assumes: Generalisation, Overfitting and Underfitting
- 0620 min
The No Free Lunch Theorem
AdvancedWhy no learner dominates across all problems, and what that means for model selection in practice.
Assumes: The Bias–Variance Trade-off
- 0734 min
VC Dimension and PAC Learning
AdvancedShattering, VC dimension, sample complexity bounds, and the theory behind how much data is enough.
Assumes: Empirical Risk Minimisation · Probability Inequalities
- 0822 min
Train, Validation and Test Splits
BeginnerThe role of each split, why the test set must stay untouched, and stratification.
Assumes: Generalisation, Overfitting and Underfitting
- 0930 min
Cross-Validation
Intermediatek-fold, stratified, leave-one-out and nested CV, with the bias–variance trade-off in choosing k.
Assumes: Train, Validation and Test Splits
- 1028 min
Hyperparameter Search
IntermediateGrid, random and Bayesian optimisation, successive halving, and budgeting search honestly.
Assumes: Cross-Validation · Bayesian Optimisation
- 1130 min
Classification Metrics
BeginnerConfusion matrix, accuracy, precision, recall, F1, specificity and Cohen's kappa, all computed by hand.
Assumes: Train, Validation and Test Splits
- 1228 min
ROC and Precision–Recall Curves
IntermediateThreshold sweeps, AUC interpretation, and why PR curves beat ROC under heavy imbalance.
Assumes: Classification Metrics
- 1324 min
Regression Metrics
BeginnerMSE, RMSE, MAE, MAPE, R² and adjusted R², and which to report for which audience.
Assumes: Train, Validation and Test Splits
- 1426 min
Probability Calibration
AdvancedReliability diagrams, Brier score, Platt scaling and isotonic regression.
Assumes: ROC and Precision–Recall Curves
- 1528 min
Handling Class Imbalance
IntermediateResampling, SMOTE, class weights, threshold tuning, and choosing metrics that survive skew.
Assumes: ROC and Precision–Recall Curves
- 1630 min
Feature Engineering
IntermediateTransformations, interactions, binning, domain features, and why this still outperforms model tinkering.
Assumes: Encoding Categorical Features
- 1728 min
Feature Selection
IntermediateFilter, wrapper and embedded methods; mutual information, RFE and stability selection.
Assumes: Feature Engineering
- 1830 min
Regularisation
IntermediateL1 and L2 penalties, elastic net, the constrained-optimisation view, and why L1 induces sparsity.
Assumes: The Bias–Variance Trade-off · Lagrange Multipliers
- 1926 min
The Curse of Dimensionality
IntermediateVolume concentration, distance concentration, sample-density collapse, and its consequences for kNN and kernels.
Assumes: The Bias–Variance Trade-off
- 2028 min
Pipelines and Data Leakage
IntermediateThe many ways leakage sneaks in — scaling before splitting, target encoding, temporal leaks — and how pipelines prevent it.
Assumes: Cross-Validation