MODULE 16
Reinforcement Learning
Bandits and MDPs through to PPO: the full progression from tabular dynamic programming to deep policy-gradient methods.
26 lessons~13h reading
- 0124 min
The Reinforcement Learning Problem
BeginnerComing soonAgent, environment, reward and the exploration–exploitation dilemma, contrasted with supervised learning.
- 0228 min
Multi-Armed Bandits
IntermediateComing soonThe simplest RL problem: action values, regret, and incremental mean updates.
Assumes: The Reinforcement Learning Problem
- 0330 min
Exploration Strategies
AdvancedComing soonEpsilon-greedy, optimistic initialisation, UCB and Thompson sampling, compared on regret.
Assumes: Multi-Armed Bandits
- 0430 min
Markov Decision Processes
IntermediateComing soonStates, actions, transitions, rewards and the Markov property; the MDP as the formal RL object.
Assumes: The Reinforcement Learning Problem · Markov Chains
- 0528 min
Returns, Policies and Value Functions
IntermediateComing soonDiscounted return, stochastic policies, and the state- and action-value functions.
Assumes: Markov Decision Processes
- 0632 min
The Bellman Equations
AdvancedComing soonExpectation and optimality equations derived from first principles, in scalar and matrix form.
Assumes: Returns, Policies and Value Functions
- 0728 min
Iterative Policy Evaluation
IntermediateComing soonComputing v_pi by repeated sweeps, with a full numeric trace on a grid world.
Assumes: The Bellman Equations
- 0828 min
Policy Iteration
IntermediateComing soonAlternating evaluation and greedy improvement, the policy improvement theorem, and convergence.
Assumes: Iterative Policy Evaluation
- 0928 min
Value Iteration
IntermediateComing soonFolding improvement into the update, the contraction-mapping argument, and stopping criteria.
Assumes: Policy Iteration
- 1030 min
Monte Carlo Methods
IntermediateComing soonLearning from complete episodes, first-visit vs every-visit, and off-policy importance sampling.
Assumes: Value Iteration
- 1130 min
Temporal Difference Learning
AdvancedComing soonTD(0), bootstrapping, the TD error, and the bias–variance comparison with Monte Carlo.
Assumes: Monte Carlo Methods
- 1226 min
SARSA
AdvancedComing soonOn-policy TD control, the quintuple update, and the cliff-walking comparison with Q-learning.
Assumes: Temporal Difference Learning
- 1332 min
Q-Learning
AdvancedComing soonOff-policy control, the max operator, convergence conditions, and a Q-table updated step by step.
Assumes: SARSA
- 1424 min
Expected SARSA and Double Q-Learning
AdvancedComing soonReducing update variance, maximisation bias, and the double estimator fix.
Assumes: Q-Learning
- 1530 min
n-Step Methods and Eligibility Traces
AdvancedComing soonInterpolating between TD and Monte Carlo, the lambda-return, and TD(lambda) with traces.
Assumes: Q-Learning
- 1630 min
Value Function Approximation
AdvancedComing soonReplacing tables with parameterised functions, the deadly triad, and semi-gradient methods.
Assumes: n-Step Methods and Eligibility Traces
- 1732 min
Deep Q-Networks
AdvancedComing soonNeural Q-functions, experience replay, target networks, and the Atari result.
Assumes: Value Function Approximation · Adaptive Optimisers: AdaGrad to AdamW
- 1828 min
Rainbow: DQN Improvements
AdvancedComing soonDouble DQN, duelling architectures, prioritised replay, noisy nets and distributional RL.
Assumes: Deep Q-Networks
- 1934 min
The Policy Gradient Theorem
AdvancedComing soonParameterising policies directly, the log-derivative trick, and the full theorem derivation.
Assumes: Value Function Approximation
- 2028 min
REINFORCE
AdvancedComing soonMonte Carlo policy gradient, baselines for variance reduction, and the full algorithm.
Assumes: The Policy Gradient Theorem
- 2130 min
Actor–Critic Methods
AdvancedComing soonCombining learned value estimates with policy gradients, and the advantage function.
Assumes: REINFORCE
- 2224 min
A2C and A3C
AdvancedComing soonSynchronous and asynchronous parallel actor-critics, and why parallelism stabilises training.
Assumes: Actor–Critic Methods
- 2334 min
TRPO and PPO
AdvancedComing soonTrust regions, the surrogate objective, the clipped ratio, and why PPO became the RLHF default.
Assumes: Actor–Critic Methods
- 2430 min
DDPG, TD3 and SAC
AdvancedComing soonContinuous action spaces, deterministic policy gradients, and maximum-entropy RL.
Assumes: TRPO and PPO
- 2530 min
Model-Based Reinforcement Learning
AdvancedComing soonLearning environment models, Dyna, planning with learned models, and MCTS.
Assumes: Value Iteration
- 2628 min
Offline Reinforcement Learning
AdvancedComing soonLearning from fixed datasets, distributional shift, conservative Q-learning and behaviour cloning.
Assumes: TRPO and PPO