# Extraction Patterns and "Unknown" Handling — Prompt Engineering

Source: https://www.geekswithgeeks.com/en/prompt-engineering/s-extract

> Pull fields from text without inventing missing values.

## Do not let it guess

Extraction (names, dates, amounts, entities) is one of the most reliable LLM uses if you control the failure case. Tell the model to use **only the text**, to return `null` or `"unknown"` for absent fields, to copy values **exactly as written** (or in a specified normalised form such as ISO dates), and optionally to return the **quote** that supports each value so you can verify it appears in the source. Add one example with a missing field. Then check in code that each returned quote is really a substring of the input.

## An extraction prompt

Asks for evidence quotes and explicit nulls. Illustrative; not run here.

```text
Extract the fields from <email>. Use ONLY the email text.
Return JSON: {"order_id": str|null, "amount": number|null, "quote_order": str|null, "quote_amount": str|null}
If a field is not stated, use null. Each quote must be copied exactly from the email.

<email>
Hi, my order 4821 was delayed. Please advise.
</email>

Expected: {"order_id": "4821", "amount": null, "quote_order": "order 4821", "quote_amount": null}
```

## Normalise formats you control

Ask for ISO dates (2026-10-02) and plain numbers without currency symbols, and parse them in code.

**Quiz:** Why ask for supporting quotes?

- [ ] To skip validation
- [ ] To make the reply longer
- [x] Code can verify each value appears in the source
- [ ] To change the model

*Answer:* Code can verify each value appears in the source. A quote can be checked mechanically, catching invented values.
