MODULE 17
Natural Language Processing
Tokenisation and n-grams through embeddings, attention and the complete transformer, with shapes traced end to end.
32 lessons~15h reading
- 0122 min
The NLP Pipeline
BeginnerComing soonWhy language is hard for machines: ambiguity, compositionality and the classical processing stages.
- 0222 min
Text Normalisation
BeginnerComing soonCase folding, Unicode normalisation, accent stripping, whitespace and punctuation handling.
Assumes: The NLP Pipeline
- 0326 min
Regular Expressions for Text
BeginnerComing soonCharacter classes, quantifiers, groups, lookarounds, and catastrophic backtracking.
Assumes: Text Normalisation
- 0422 min
Stemming and Lemmatisation
BeginnerComing soonPorter and Snowball stemmers versus dictionary lemmatisation, and the precision trade-off.
Assumes: Text Normalisation
- 0530 min
n-Gram Language Models
IntermediateComing soonThe Markov assumption over text, maximum likelihood estimates, and generating from an n-gram model.
Assumes: Markov Chains · Stemming and Lemmatisation
- 0630 min
Smoothing
AdvancedComing soonLaplace, add-k, Good–Turing, backoff, interpolation and Kneser–Ney, worked on a small corpus.
Assumes: n-Gram Language Models
- 0726 min
Perplexity
IntermediateComing soonCross-entropy and perplexity derived, computed by hand, and its pitfalls as a comparison metric.
Assumes: Smoothing
- 0824 min
Bag of Words and Count Vectors
BeginnerComing soonDocument–term matrices, vocabulary construction, sparsity, and the loss of word order.
Assumes: Stemming and Lemmatisation
- 0928 min
TF-IDF
BeginnerComing soonTerm frequency and inverse document frequency variants, the full formula derived and computed by hand.
Assumes: Bag of Words and Count Vectors
- 1026 min
Vector Space Retrieval
IntermediateComing soonCosine similarity ranking, length normalisation, and precision/recall for retrieval.
Assumes: TF-IDF
- 1132 min
word2vec: CBOW and Skip-Gram
AdvancedComing soonDistributional semantics, both architectures, and the embedding geometry that makes analogies work.
Assumes: Vector Space Retrieval · Backpropagation
- 1230 min
Negative Sampling and Hierarchical Softmax
AdvancedComing soonWhy the full softmax is intractable, the negative sampling objective derived, and subsampling frequent words.
Assumes: word2vec: CBOW and Skip-Gram
- 1326 min
GloVe and fastText
AdvancedComing soonGlobal co-occurrence factorisation, and subword embeddings that handle unseen words.
Assumes: Negative Sampling and Hierarchical Softmax
- 1428 min
Embedding Geometry and Bias
AdvancedComing soonSimilarity, analogy arithmetic, anisotropy, and measuring and mitigating social bias in embeddings.
Assumes: GloVe and fastText
- 1526 min
Part-of-Speech Tagging
IntermediateComing soonTagsets, ambiguity, rule-based and statistical tagging, and HMM taggers.
Assumes: Hidden Markov Models
- 1630 min
Sequence Labelling and CRFs
AdvancedComing soonThe BIO scheme, linear-chain conditional random fields, and why CRFs beat independent classification.
Assumes: Part-of-Speech Tagging
- 1726 min
Named Entity Recognition
IntermediateComing soonEntity types, span-based evaluation, nested entities, and modern neural NER.
Assumes: Sequence Labelling and CRFs
- 1830 min
Syntactic Parsing
AdvancedComing soonConstituency and dependency grammars, CKY parsing, transition-based parsing and treebanks.
Assumes: Named Entity Recognition
- 1928 min
Neural Machine Translation
AdvancedComing soonEncoder–decoder translation, the bottleneck problem, and what motivated attention.
Assumes: Sequence-to-Sequence Models
- 2034 min
The Attention Mechanism
AdvancedComing soonQueries, keys and values; additive vs dot-product attention, and an attention matrix computed by hand.
Assumes: Neural Machine Translation
- 2138 min
The Transformer Architecture
AdvancedComing soonThe full encoder–decoder stack: attention sublayers, feed-forward blocks, residuals and normalisation.
Assumes: The Attention Mechanism · Layer, Group and RMS Normalisation
- 2232 min
Multi-Head and Masked Attention
AdvancedComing soonSplitting into heads, scaled dot-product attention, causal masking and padding masks.
Assumes: The Transformer Architecture
- 2328 min
Positional Encoding
AdvancedComing soonWhy attention is permutation-invariant, sinusoidal encodings derived, and learned alternatives.
Assumes: Multi-Head and Masked Attention
- 2434 min
Transformer Shapes: End-to-End Walkthrough
AdvancedComing soonEvery tensor shape from token ids to logits for a concrete small model, with parameter counts.
Assumes: Positional Encoding
- 2532 min
BERT and Masked Language Modelling
AdvancedComing soonBidirectional pretraining, the MLM and NSP objectives, and fine-tuning for downstream tasks.
Assumes: Transformer Shapes: End-to-End Walkthrough
- 2626 min
Encoder, Decoder and Encoder–Decoder Families
IntermediateComing soonBERT vs GPT vs T5: which architecture suits which task, and why decoder-only won for generation.
Assumes: BERT and Masked Language Modelling
- 2732 min
Subword Tokenisation
AdvancedComing soonBPE merges traced by hand, WordPiece, Unigram and SentencePiece, plus vocabulary-size trade-offs.
Assumes: Bag of Words and Count Vectors
- 2826 min
Text Classification
IntermediateComing soonClassical baselines through fine-tuned transformers, with class imbalance and multi-label handling.
Assumes: BERT and Masked Language Modelling
- 2932 min
Topic Modelling
AdvancedComing soonLSA via SVD, probabilistic LSA, LDA with its generative story, coherence metrics and BERTopic.
Assumes: TF-IDF · PCA via SVD
- 3024 min
Summarisation
IntermediateComing soonExtractive versus abstractive approaches, and the faithfulness problem in generated summaries.
Assumes: Encoder, Decoder and Encoder–Decoder Families
- 3126 min
Question Answering
IntermediateComing soonExtractive span prediction, open-domain QA, and the retrieval bridge to RAG.
Assumes: BERT and Masked Language Modelling
- 3230 min
NLP Evaluation Metrics
IntermediateComing soonBLEU and ROUGE computed by hand, plus METEOR, chrF, BERTScore and their known weaknesses.
Assumes: Summarisation