# Revision: Cheat Sheet and Self-Check — Claude API / OpenAI API Basics

Source: https://www.geekswithgeeks.com/en/llm-apis/z-revision

> Review the key ideas of the whole course.

## Cheat sheet

**Basics**: HTTPS + JSON, stateless (resend history), key in env var or secrets manager, never in client code; Anthropic `/v1/messages` (`x-api-key`, `anthropic-version`, top-level `system`, required `max_tokens`, `content[0].text`, `stop_reason`) versus OpenAI Chat Completions (`Authorization: Bearer`, `system` message, `choices[0].message.content`, `finish_reason`); newer OpenAI Responses API exists; model names change, keep them in config. **Parameters**: `max_tokens` (check truncation), temperature low for extraction, one of temperature/top_p, stop sequences. **Output**: ask for JSON, extract, validate, retry with the error. **Streaming**: SSE events, assemble text, cancel upstream, relay via backend. **Tools**: model requests, YOUR code validates and runs, return `tool_result` / `role: "tool"` messages, cap loop, least privilege, approvals. **Reliability**: retry 429/5xx/timeouts with exponential backoff + jitter and a cap, never retry 400/401, honour retry-after, limit concurrency, timeouts, fallbacks, circuit breaker. **Cost**: input/output tokens priced separately; history, system prompt and tool definitions count every call; trim, cap output, smaller model, prompt caching, batch; track usage per feature. **Security**: redact logs, escape output, injection defence, data minimisation. **Testing**: fake server for integration; separate eval set for quality.

**Quiz:** Your call returns 429 Too Many Requests. What is the right response?

- [ ] Give up and delete the data
- [ ] Retry immediately in a tight loop
- [ ] Change the API key
- [x] Wait with exponential backoff and jitter, honour retry-after, then retry a limited number of times

*Answer:* Wait with exponential backoff and jitter, honour retry-after, then retry a limited number of times. 429 is transient; polite, bounded retries usually succeed.

**Quiz:** A reply stopped with stop_reason "max_tokens". What does this mean?

- [ ] The model finished perfectly
- [x] The output was truncated at your limit and may be incomplete
- [ ] The key expired
- [ ] A tool was requested

*Answer:* The output was truncated at your limit and may be incomplete. Truncation must be handled: raise the limit, shorten the request or continue.

**Quiz:** Which is the safest place to enforce "a user may only see their own orders" in a tool-using assistant?

- [ ] In the system prompt only
- [x] In the tool code, filtering by the authenticated user id
- [ ] In the model's temperature
- [ ] In the browser

*Answer:* In the tool code, filtering by the authenticated user id. Server-side code enforcement cannot be bypassed by clever prompting.
