# Structured Output, Validation and Repair Loops — LLM Engineering Foundations

Source: https://www.geekswithgeeks.com/en/llm-engineering/q-structured

> Make model output safe for code to consume.

## Contract, validate, repair, fall back

Downstream code needs predictable data, not prose. Define the **contract** (a schema with fields, types and allowed values) and request it using the provider's structured-output or tool-calling features where available. Then **validate in code**: parse, check types, ranges and enumerations, and cross-field rules (for example, "refund amount must not exceed order total"). On failure, run a **bounded repair loop**: re-ask once or twice with the specific validation error appended; if it still fails, **fall back safely** (a default, a simpler path, or a human queue) and count the failure. Track the **valid-output rate** as a core metric: a drop is an early warning of prompt or model drift. Never let unvalidated model output reach a database, payment system or user-visible action.

## A validate-and-repair loop (illustrative)

`call_model` and `fallback_to_human` are placeholders for your own functions. Not run here.

```python
import json

ALLOWED = {"billing", "technical", "other"}

def validate(raw):
    data = json.loads(raw)                                   # may raise: not JSON
    if data.get("category") not in ALLOWED: raise ValueError("category must be one of %s" % sorted(ALLOWED))
    if not (1 <= int(data.get("urgency", 0)) <= 5): raise ValueError("urgency must be 1..5")
    return data

def classify(ticket, max_repairs=2):
    prompt = f"Classify this ticket as JSON: {ticket}"
    for attempt in range(1 + max_repairs):
        raw = call_model(prompt)
        try:
            return validate(raw)
        except (ValueError, json.JSONDecodeError) as e:
            prompt += f"\nYour last reply was invalid ({e}). Reply with valid JSON only."
    return fallback_to_human(ticket)                           # bounded: never loop forever
```

## Count repairs as a metric

A rising repair rate warns of prompt or model drift before users notice.

**Quiz:** Why bound the repair loop?

- [ ] Because models dislike retries
- [x] To avoid endless retries and runaway cost; fall back safely instead
- [ ] To make outputs longer
- [ ] Repair loops should be unlimited

*Answer:* To avoid endless retries and runaway cost; fall back safely instead. Every retry costs money and time; cap them and have a safe default.
