Large Language Models

Understand how LLMs work from tokens and embeddings to attention, training, prompting, evaluation, safety and deployment, with small runnable models you can verify.

Start course →

Syllabus

What an LLM Is

  1. What Is a Large Language Model
  2. The Life of a Model: Pretraining to Chat
  3. Tokens and Byte-Pair Encoding
  4. What LLMs Are Good and Bad At

From Text to Probabilities

  1. Embeddings and Similarity
  2. Logits, Softmax and Temperature
  3. Sampling: Greedy, Top-k and Top-p
  4. A Tiny Language Model You Can Run
  5. Loss and Perplexity

The Transformer

  1. Self-Attention: Queries, Keys, Values
  2. Causal Masking and Generation
  3. Layers, Feed-Forward Blocks and Positions
  4. Context Window and the KV Cache

Training and Adapting Models

  1. Scaling Laws and Data
  2. Fine-Tuning, LoRA and When Not to Fine-Tune
  3. Alignment: RLHF and Preference Optimisation

Building with LLMs

  1. Prompting Basics
  2. Retrieval-Augmented Generation (RAG)
  3. Tool Use and Structured Output
  4. Cost, Latency and Caching

Evaluation and Safety

  1. Evaluating LLM Output
  2. Reducing Hallucination
  3. Prompt Injection, Privacy and Bias

Running Models in Production

  1. Model Size and Quantization
  2. Serving, Batching and Choosing a Model

Putting It Together

  1. Case Study: A Policy Question-Answering Assistant
  2. Revision: Cheat Sheet and Self-Check