# From Prototype to Production: A Staged Path — LLM Engineering Foundations

Source: https://www.geekswithgeeks.com/en/llm-engineering/i-path

> Move through stages, each with entry criteria.

## Earn each stage

A sensible path: (1) **Prototype**: prove the idea works on a handful of examples with a strong model; no commitments. (2) **Evaluation set**: collect 50 to 200 realistic cases and a scoring method; now you have a baseline. (3) **Iterate**: improve prompts, retrieval, model choice, and measure each change against the set; set the cheapest configuration that meets the target. (4) **Harden**: add validation, retries, timeouts, safety controls, access control, logging and cost caps. (5) **Pilot**: release to a small group or shadow mode (run alongside humans without acting) and compare with real outcomes. (6) **Roll out gradually** with monitoring and a rollback switch. (7) **Operate**: watch quality, cost and drift; fold failures back into the test set. Each stage has **entry criteria** (for example, "target quality reached on the offline set") so you do not skip ahead. Many projects stall at stage 1 because they never build the evaluation set; that is the single most valuable early investment.

## The stages and their entry criteria

Do not skip a stage without a reason.

```text
Stage          Entry criterion                                  Output
prototype      an idea and a few examples                       "it can work"
eval set       prototype shows promise                          50-200 cases + scorer + baseline
iterate        eval set exists                                  config that hits the quality target at acceptable cost
harden         quality target met offline                       validation, retries, ACLs, caps, logging
pilot/shadow   hardening checklist done; red-team suite passes  real-world comparison with humans
rollout        pilot meets targets                              gradual traffic, alerts, rollback switch
operate        in production                                    dashboards; failures feed the eval set
```

## Shadow before users

Running beside humans without acting is the safest way to learn how the system behaves on real inputs.

**Quiz:** What is the most valuable early investment in an LLM project?

- [ ] The longest possible system prompt
- [x] A realistic evaluation set with a scoring method
- [ ] The largest model available
- [ ] A fancy logo

*Answer:* A realistic evaluation set with a scoring method. Everything else depends on being able to measure.
