Lesson 13 / 25

Structured Output and Extraction

Get reliable JSON from a model using the Information Extractor or an output parser.

Make the model fill a form

Free-text answers are hard to use in later nodes. The Information Extractor node takes text and a list of fields you want (name, date, amount) and returns them as structured data. An output parser with a JSON schema on an LLM chain does the same job. Always validate the result, because the model may return null for missing fields or misread a number.

A schema for an invoice

Descriptions on each property guide the model. Dates in ISO format are easy for later nodes to parse.

{
  "type": "object",
  "properties": {
    "vendor": { "type": "string", "description": "Company that issued the invoice" },
    "date":   { "type": "string", "description": "Invoice date in YYYY-MM-DD" },
    "total":  { "type": "number", "description": "Total amount payable, numeric only" }
  },
  "required": ["vendor", "total"]
}

Validate in a Code node

After extraction, add a small Code or IF node that checks the total is a positive number and the date parses. Route failures to a human-review path instead of posting bad data to your accounting system.

Quick check: Why validate data after an Information Extractor step?

  • Extractors never work
  • The model might misread values or return null
  • Validation makes data longer
  • n8n deletes unvalidated data
Answer

The model might misread values or return null — A well-formed JSON object can still contain a wrong number or missing value.