Lesson 2 / 29
A Reference Architecture for an LLM Feature
Name the components that surround the model call.
The model call is the small part
Look inside a production LLM feature and the model call is surrounded by ordinary software. Input handling: authentication, rate limits, size limits, sanitising, language detection. Context building: prompt templates, retrieval (RAG), conversation memory, user and tenant data with access control. Model layer: provider SDK or self-hosted server, model choice or routing, timeouts, retries, fallbacks, caching. Output handling: parsing, schema validation, repair or retry, safety checks, formatting, citations. Tools and actions: policy engine, approvals, audit log. Observability: traces, metrics, cost tracking, user feedback. Evaluation and release: test sets, CI gates, prompt registry, staged rollout. Drawing this on one page for your feature shows where each earlier course fits, and where the gaps are.
The pieces around the model call
Read left to right along the request path.
request -> [auth, rate limit, size limit] -> [build context: template + retrieval + memory + ACL]
-> [model layer: route, cache, timeout, retry, fallback] -> [parse + validate + repair]
-> [safety checks, policy for tool calls, approvals] -> response
alongside: traces + metrics + cost | feedback | eval sets + CI gate | prompt/model registry | staged rolloutDraw it before you build it
A one-page diagram of entry points, model layer and exits exposes missing controls early.
Quick check: Which component sits after the model call and before the user sees a response?
- Parsing, validation and safety checks
- Training the base model
- Choosing the GPU
- Writing the system prompt
Answer
Parsing, validation and safety checks — Output handling turns raw text into something safe and usable.