Lesson 1 / 27

What Is a Large Language Model

Define an LLM and what it actually does at each step.

A next-token predictor

A large language model (LLM) is a neural network with billions of numerical parameters, trained on huge amounts of text to do one job: given the text so far, produce a probability for every possible next token. A program then picks one token (by rule or by random draw), appends it and repeats. Chat, translation, summaries and code all come from this single loop. The model does not look things up in a database of facts; whatever it "knows" is stored implicitly in its parameters, which is why it can be fluent and still wrong.

Predict the next token

An LLM is a very large neural network trained to predict the next piece of text.

Four steps: text, tokens, probabilities, next token.
Figure 1.1 — Text, tokens, probabilities and the next token.

Phone autocomplete, scaled up

Your phone suggests the next word from what you typed. An LLM is the same idea with a vastly larger memory of language and a much longer view of the text so far.

Parameters are learned numbers

"7B" means seven billion learned numbers. More parameters usually mean more capacity but also more memory and cost.

Quick check: What is the basic task an LLM is trained for?

  • Run SQL queries
  • Look up answers in a fact database
  • Predict the next token given the previous ones
  • Compress images
Answer

Predict the next token given the previous ones — Everything an LLM produces is built by repeated next-token prediction.