Skip to content
VibeFormer

MODULE 17

Natural Language Processing

Tokenisation and n-grams through embeddings, attention and the complete transformer, with shapes traced end to end.

32 lessons~15h reading

  1. 01

    The NLP Pipeline

    BeginnerComing soon

    Why language is hard for machines: ambiguity, compositionality and the classical processing stages.

    22 min
  2. 02

    Text Normalisation

    BeginnerComing soon

    Case folding, Unicode normalisation, accent stripping, whitespace and punctuation handling.

    Assumes: The NLP Pipeline

    22 min
  3. 03

    Regular Expressions for Text

    BeginnerComing soon

    Character classes, quantifiers, groups, lookarounds, and catastrophic backtracking.

    Assumes: Text Normalisation

    26 min
  4. 04

    Stemming and Lemmatisation

    BeginnerComing soon

    Porter and Snowball stemmers versus dictionary lemmatisation, and the precision trade-off.

    Assumes: Text Normalisation

    22 min
  5. 05

    n-Gram Language Models

    IntermediateComing soon

    The Markov assumption over text, maximum likelihood estimates, and generating from an n-gram model.

    Assumes: Markov Chains · Stemming and Lemmatisation

    30 min
  6. 06

    Smoothing

    AdvancedComing soon

    Laplace, add-k, Good–Turing, backoff, interpolation and Kneser–Ney, worked on a small corpus.

    Assumes: n-Gram Language Models

    30 min
  7. 07

    Perplexity

    IntermediateComing soon

    Cross-entropy and perplexity derived, computed by hand, and its pitfalls as a comparison metric.

    Assumes: Smoothing

    26 min
  8. 08

    Bag of Words and Count Vectors

    BeginnerComing soon

    Document–term matrices, vocabulary construction, sparsity, and the loss of word order.

    Assumes: Stemming and Lemmatisation

    24 min
  9. 09

    TF-IDF

    BeginnerComing soon

    Term frequency and inverse document frequency variants, the full formula derived and computed by hand.

    Assumes: Bag of Words and Count Vectors

    28 min
  10. 10

    Vector Space Retrieval

    IntermediateComing soon

    Cosine similarity ranking, length normalisation, and precision/recall for retrieval.

    Assumes: TF-IDF

    26 min
  11. 11

    word2vec: CBOW and Skip-Gram

    AdvancedComing soon

    Distributional semantics, both architectures, and the embedding geometry that makes analogies work.

    Assumes: Vector Space Retrieval · Backpropagation

    32 min
  12. 12

    Negative Sampling and Hierarchical Softmax

    AdvancedComing soon

    Why the full softmax is intractable, the negative sampling objective derived, and subsampling frequent words.

    Assumes: word2vec: CBOW and Skip-Gram

    30 min
  13. 13

    GloVe and fastText

    AdvancedComing soon

    Global co-occurrence factorisation, and subword embeddings that handle unseen words.

    Assumes: Negative Sampling and Hierarchical Softmax

    26 min
  14. 14

    Embedding Geometry and Bias

    AdvancedComing soon

    Similarity, analogy arithmetic, anisotropy, and measuring and mitigating social bias in embeddings.

    Assumes: GloVe and fastText

    28 min
  15. 15

    Part-of-Speech Tagging

    IntermediateComing soon

    Tagsets, ambiguity, rule-based and statistical tagging, and HMM taggers.

    Assumes: Hidden Markov Models

    26 min
  16. 16

    Sequence Labelling and CRFs

    AdvancedComing soon

    The BIO scheme, linear-chain conditional random fields, and why CRFs beat independent classification.

    Assumes: Part-of-Speech Tagging

    30 min
  17. 17

    Named Entity Recognition

    IntermediateComing soon

    Entity types, span-based evaluation, nested entities, and modern neural NER.

    Assumes: Sequence Labelling and CRFs

    26 min
  18. 18

    Syntactic Parsing

    AdvancedComing soon

    Constituency and dependency grammars, CKY parsing, transition-based parsing and treebanks.

    Assumes: Named Entity Recognition

    30 min
  19. 19

    Neural Machine Translation

    AdvancedComing soon

    Encoder–decoder translation, the bottleneck problem, and what motivated attention.

    Assumes: Sequence-to-Sequence Models

    28 min
  20. 20

    The Attention Mechanism

    AdvancedComing soon

    Queries, keys and values; additive vs dot-product attention, and an attention matrix computed by hand.

    Assumes: Neural Machine Translation

    34 min
  21. 21

    The Transformer Architecture

    AdvancedComing soon

    The full encoder–decoder stack: attention sublayers, feed-forward blocks, residuals and normalisation.

    Assumes: The Attention Mechanism · Layer, Group and RMS Normalisation

    38 min
  22. 22

    Multi-Head and Masked Attention

    AdvancedComing soon

    Splitting into heads, scaled dot-product attention, causal masking and padding masks.

    Assumes: The Transformer Architecture

    32 min
  23. 23

    Positional Encoding

    AdvancedComing soon

    Why attention is permutation-invariant, sinusoidal encodings derived, and learned alternatives.

    Assumes: Multi-Head and Masked Attention

    28 min
  24. 24

    Transformer Shapes: End-to-End Walkthrough

    AdvancedComing soon

    Every tensor shape from token ids to logits for a concrete small model, with parameter counts.

    Assumes: Positional Encoding

    34 min
  25. 25

    BERT and Masked Language Modelling

    AdvancedComing soon

    Bidirectional pretraining, the MLM and NSP objectives, and fine-tuning for downstream tasks.

    Assumes: Transformer Shapes: End-to-End Walkthrough

    32 min
  26. 26

    Encoder, Decoder and Encoder–Decoder Families

    IntermediateComing soon

    BERT vs GPT vs T5: which architecture suits which task, and why decoder-only won for generation.

    Assumes: BERT and Masked Language Modelling

    26 min
  27. 27

    Subword Tokenisation

    AdvancedComing soon

    BPE merges traced by hand, WordPiece, Unigram and SentencePiece, plus vocabulary-size trade-offs.

    Assumes: Bag of Words and Count Vectors

    32 min
  28. 28

    Text Classification

    IntermediateComing soon

    Classical baselines through fine-tuned transformers, with class imbalance and multi-label handling.

    Assumes: BERT and Masked Language Modelling

    26 min
  29. 29

    Topic Modelling

    AdvancedComing soon

    LSA via SVD, probabilistic LSA, LDA with its generative story, coherence metrics and BERTopic.

    Assumes: TF-IDF · PCA via SVD

    32 min
  30. 30

    Summarisation

    IntermediateComing soon

    Extractive versus abstractive approaches, and the faithfulness problem in generated summaries.

    Assumes: Encoder, Decoder and Encoder–Decoder Families

    24 min
  31. 31

    Question Answering

    IntermediateComing soon

    Extractive span prediction, open-domain QA, and the retrieval bridge to RAG.

    Assumes: BERT and Masked Language Modelling

    26 min
  32. 32

    NLP Evaluation Metrics

    IntermediateComing soon

    BLEU and ROUGE computed by hand, plus METEOR, chrF, BERTScore and their known weaknesses.

    Assumes: Summarisation

    30 min