Lesson 15 / 27
Error Types and Which to Retry
Map status codes to actions.
Retry the transient, fix the permanent
HTTP status codes tell you what to do. 400 invalid request (a bug in your request, such as a missing max_tokens or a malformed message list): do not retry, fix the code. 401/403 authentication or permission problem: do not retry, check the key and access. 404 unknown model or endpoint. 413 request too large. 429 rate limited (too many requests or tokens per minute) and 5xx server errors or overloaded: transient, retry after waiting (honour a retry-after header if present). Network timeouts and dropped connections are also retryable. The official SDKs already retry some of these automatically (a small default number of retries with backoff) and raise typed exceptions such as RateLimitError, BadRequestError and APIConnectionError, each carrying the status code and the error body.
Expect failure, recover gracefully
Networks and rate limits fail; classify errors, retry the right ones with backoff, and degrade gracefully.
Retries and typed errors, run
I ran this in a Python virtual environment with the anthropic 1.11.0 and openai 3.22.1 SDKs against a small local stand-in server (shown in the testing topic, saved as mock.py). The server returns canned replies, so no key, network or real model is involved: it proves how the SDK builds requests and handles replies, not what a real model would say. The stand-in server answers 429 once. With max_retries=2 the SDK retries by itself and succeeds on the second request. With max_retries=0 the same 429 raises RateLimitError after one request. A missing max_tokens (set to 0) returns a 400 that raises BadRequestError with the server's message, and the OpenAI SDK raises its own RateLimitError for a 429.
import mock, anthropic, openai
srv = mock.start(); base = f"http://127.0.0.1:{srv.server_address[1]}" # local stand-in server, not a real API
body = dict(model="demo-model", max_tokens=50, messages=[{"role": "user", "content": "Hi"}])
mock.STATE["log"].clear(); mock.STATE["fail_next"] = 1 # server answers 429 once
ok = anthropic.Anthropic(api_key="k", base_url=base, max_retries=2).messages.create(**body)
print("with retries=2 : succeeded after", len(mock.STATE["log"]), "requests ->", ok.content[0].text[:5])
mock.STATE["log"].clear(); mock.STATE["fail_next"] = 1
try:
anthropic.Anthropic(api_key="k", base_url=base, max_retries=0).messages.create(**body)
except anthropic.RateLimitError as e:
print("with retries=0 : RateLimitError", e.status_code, "after", len(mock.STATE["log"]), "request")
try:
anthropic.Anthropic(api_key="k", base_url=base).messages.create(**{**body, "max_tokens": 0})
except anthropic.BadRequestError as e:
print("bad request :", e.status_code, "-", e.body["error"]["message"])
mock.STATE["fail_next"] = 1
try:
openai.OpenAI(api_key="k", base_url=base + "/v1", max_retries=0).chat.completions.create(
model="demo-model", messages=[{"role": "user", "content": "Hi"}])
except openai.RateLimitError as e:
print("openai 429 :", type(e).__name__, e.status_code)
Output:
with retries=2 : succeeded after 2 requests -> Paris with retries=0 : RateLimitError 429 after 1 request bad request : 400 - max_tokens: Field required openai 429 : RateLimitError 429
Log the request id
Providers return a request id in response headers or the error body. Log it so support can trace a failed call.
Quick check: Which error should you NOT retry?
- 400 invalid request
- 429 rate limited
- 503 overloaded
- A dropped connection
Answer
400 invalid request — Repeating a malformed request just fails again; fix the request.