Lesson 17 / 27

Rate Limits, Concurrency and Timeouts

Stay within quotas and bound waiting time.

Quotas on requests and on tokens

Providers limit requests per minute and tokens per minute (separately for input and output, by model and account tier), and sometimes concurrent requests. Exceeding them gives 429. Good clients: read the rate-limit headers to see remaining quota, queue work and cap concurrency (a worker pool rather than firing 1,000 requests at once), spread batch jobs over time, and use the provider's batch interface for non-urgent bulk work (usually cheaper). Set timeouts on every call (the SDKs have configurable defaults, and generation can legitimately take a while, especially for long outputs), and also an overall deadline for the user-facing request so slow calls cannot pile up and exhaust your threads.

A worker-pool pattern (illustrative)

Bounded concurrency keeps you under the rate limit. Not run here because it would call a real API.

from concurrent.futures import ThreadPoolExecutor

def summarise(doc):
    return client.messages.create(model=MODEL, max_tokens=200, timeout=30,
                                  messages=[{"role": "user", "content": f"Summarise: {doc}"}])

with ThreadPoolExecutor(max_workers=4) as pool:      # at most 4 requests in flight
    results = list(pool.map(summarise, documents))

Read the rate-limit headers

Remaining-quota headers let you slow down before you hit a 429.

Quick check: Why cap concurrency instead of sending all requests at once?

  • To make tokens free
  • Because HTTP forbids it
  • To stay under rate limits and avoid a wall of 429 errors
  • To avoid keys
Answer

To stay under rate limits and avoid a wall of 429 errors — Bounded parallelism respects quotas and smooths load.