Lesson 27 / 27

Revision: Cheat Sheet and Self-Check

Review the key ideas of the whole course.

Cheat sheet

Basics: HTTPS + JSON, stateless (resend history), key in env var or secrets manager, never in client code; Anthropic /v1/messages (x-api-key, anthropic-version, top-level system, required max_tokens, content[0].text, stop_reason) versus OpenAI Chat Completions (Authorization: Bearer, system message, choices[0].message.content, finish_reason); newer OpenAI Responses API exists; model names change, keep them in config. Parameters: max_tokens (check truncation), temperature low for extraction, one of temperature/top_p, stop sequences. Output: ask for JSON, extract, validate, retry with the error. Streaming: SSE events, assemble text, cancel upstream, relay via backend. Tools: model requests, YOUR code validates and runs, return tool_result / role: "tool" messages, cap loop, least privilege, approvals. Reliability: retry 429/5xx/timeouts with exponential backoff + jitter and a cap, never retry 400/401, honour retry-after, limit concurrency, timeouts, fallbacks, circuit breaker. Cost: input/output tokens priced separately; history, system prompt and tool definitions count every call; trim, cap output, smaller model, prompt caching, batch; track usage per feature. Security: redact logs, escape output, injection defence, data minimisation. Testing: fake server for integration; separate eval set for quality.

Quick check: Your call returns 429 Too Many Requests. What is the right response?

  • Give up and delete the data
  • Retry immediately in a tight loop
  • Change the API key
  • Wait with exponential backoff and jitter, honour retry-after, then retry a limited number of times
Answer

Wait with exponential backoff and jitter, honour retry-after, then retry a limited number of times — 429 is transient; polite, bounded retries usually succeed.

Quick check: A reply stopped with stop_reason "max_tokens". What does this mean?

  • The model finished perfectly
  • The output was truncated at your limit and may be incomplete
  • The key expired
  • A tool was requested
Answer

The output was truncated at your limit and may be incomplete — Truncation must be handled: raise the limit, shorten the request or continue.

Quick check: Which is the safest place to enforce "a user may only see their own orders" in a tool-using assistant?

  • In the system prompt only
  • In the tool code, filtering by the authenticated user id
  • In the model's temperature
  • In the browser
Answer

In the tool code, filtering by the authenticated user id — Server-side code enforcement cannot be bypassed by clever prompting.