Lesson 27 / 27
Revision: Cheat Sheet and Self-Check
Review the key ideas of the whole course.
Cheat sheet
Basics: HTTPS + JSON, stateless (resend history), key in env var or secrets manager, never in client code; Anthropic /v1/messages (x-api-key, anthropic-version, top-level system, required max_tokens, content[0].text, stop_reason) versus OpenAI Chat Completions (Authorization: Bearer, system message, choices[0].message.content, finish_reason); newer OpenAI Responses API exists; model names change, keep them in config. Parameters: max_tokens (check truncation), temperature low for extraction, one of temperature/top_p, stop sequences. Output: ask for JSON, extract, validate, retry with the error. Streaming: SSE events, assemble text, cancel upstream, relay via backend. Tools: model requests, YOUR code validates and runs, return tool_result / role: "tool" messages, cap loop, least privilege, approvals. Reliability: retry 429/5xx/timeouts with exponential backoff + jitter and a cap, never retry 400/401, honour retry-after, limit concurrency, timeouts, fallbacks, circuit breaker. Cost: input/output tokens priced separately; history, system prompt and tool definitions count every call; trim, cap output, smaller model, prompt caching, batch; track usage per feature. Security: redact logs, escape output, injection defence, data minimisation. Testing: fake server for integration; separate eval set for quality.
Quick check: Your call returns 429 Too Many Requests. What is the right response?
- Give up and delete the data
- Retry immediately in a tight loop
- Change the API key
- Wait with exponential backoff and jitter, honour retry-after, then retry a limited number of times
Answer
Wait with exponential backoff and jitter, honour retry-after, then retry a limited number of times — 429 is transient; polite, bounded retries usually succeed.
Quick check: A reply stopped with stop_reason "max_tokens". What does this mean?
- The model finished perfectly
- The output was truncated at your limit and may be incomplete
- The key expired
- A tool was requested
Answer
The output was truncated at your limit and may be incomplete — Truncation must be handled: raise the limit, shorten the request or continue.
Quick check: Which is the safest place to enforce "a user may only see their own orders" in a tool-using assistant?
- In the system prompt only
- In the tool code, filtering by the authenticated user id
- In the model's temperature
- In the browser
Answer
In the tool code, filtering by the authenticated user id — Server-side code enforcement cannot be bypassed by clever prompting.