Lesson 1 / 29

What LLM Engineering Is

See what changes when a model becomes a component in a product.

A component that does not behave like a function

A demo needs a good prompt and a few lucky examples. A product needs the model to behave acceptably across thousands of different inputs, week after week, within a budget, under attack, and while everything around it changes. LLM engineering is the discipline of making that true. It starts from what is different about the component: its output is non-deterministic (the same input can give different answers), hard to specify (there is no single correct answer to "summarise this"), costly (every call has a price and a latency), opaque (you cannot step through its reasoning) and changing (providers update or retire models). The response is familiar engineering applied with care: clear requirements and metrics, automated evaluation, version control for prompts and configuration, tests, observability, cost and latency budgets, staged releases, and safety controls. This course covers those foundations. Earlier courses cover the pieces (prompting, RAG, APIs, embeddings, security); here we connect them into a working practice.

Works on my prompt is not shipped

LLM engineering applies ordinary engineering discipline to a component that is non-deterministic, probabilistic and costly.

Four steps: define, build, measure, operate.
Figure 1.1 — Define, build, measure and operate.

Demo versus product

What changes when you ship.

Demo                              Product
5 hand-picked examples            thousands of real, messy inputs, including hostile ones
"looks good"                      measured pass rate with a confidence interval
prompt lives in a notebook        prompt + model + settings versioned and reviewed
one happy path                    retries, fallbacks, timeouts, "I don't know" paths
cost: ignored                     cost per resolved task tracked against a budget
no monitoring                     traces, dashboards, alerts, feedback loop
safety: hope                      least privilege, validation, approvals, red-team tests

Start with the failure you fear most

Ask what a wrong answer would cost, and let that set how much testing and control the feature needs.

Quick check: Which property of LLM components most changes how we test them?

  • They always run on laptops
  • They are written in Python
  • Outputs are non-deterministic and have no single correct answer
  • They never change
Answer

Outputs are non-deterministic and have no single correct answer — We therefore measure behaviour statistically over many cases.