# A Reference Architecture for an LLM Feature — LLM Engineering Foundations

Source: https://www.geekswithgeeks.com/en/llm-engineering/i-arch

> Name the components that surround the model call.

## The model call is the small part

Look inside a production LLM feature and the model call is surrounded by ordinary software. **Input handling**: authentication, rate limits, size limits, sanitising, language detection. **Context building**: prompt templates, retrieval (RAG), conversation memory, user and tenant data with access control. **Model layer**: provider SDK or self-hosted server, model choice or routing, timeouts, retries, fallbacks, caching. **Output handling**: parsing, schema validation, repair or retry, safety checks, formatting, citations. **Tools and actions**: policy engine, approvals, audit log. **Observability**: traces, metrics, cost tracking, user feedback. **Evaluation and release**: test sets, CI gates, prompt registry, staged rollout. Drawing this on one page for your feature shows where each earlier course fits, and where the gaps are.

## The pieces around the model call

Read left to right along the request path.

```text
request -> [auth, rate limit, size limit] -> [build context: template + retrieval + memory + ACL]
        -> [model layer: route, cache, timeout, retry, fallback] -> [parse + validate + repair]
        -> [safety checks, policy for tool calls, approvals] -> response

alongside:  traces + metrics + cost  |  feedback  |  eval sets + CI gate  |  prompt/model registry  |  staged rollout
```

## Draw it before you build it

A one-page diagram of entry points, model layer and exits exposes missing controls early.

**Quiz:** Which component sits after the model call and before the user sees a response?

- [x] Parsing, validation and safety checks
- [ ] Training the base model
- [ ] Choosing the GPU
- [ ] Writing the system prompt

*Answer:* Parsing, validation and safety checks. Output handling turns raw text into something safe and usable.
