# Rate Limits, Concurrency and Timeouts — Claude API / OpenAI API Basics

Source: https://www.geekswithgeeks.com/en/llm-apis/r-limits

> Stay within quotas and bound waiting time.

## Quotas on requests and on tokens

Providers limit **requests per minute** and **tokens per minute** (separately for input and output, by model and account tier), and sometimes concurrent requests. Exceeding them gives **429**. Good clients: read the **rate-limit headers** to see remaining quota, **queue** work and cap **concurrency** (a worker pool rather than firing 1,000 requests at once), spread batch jobs over time, and use the provider's **batch** interface for non-urgent bulk work (usually cheaper). Set **timeouts** on every call (the SDKs have configurable defaults, and generation can legitimately take a while, especially for long outputs), and also an overall **deadline** for the user-facing request so slow calls cannot pile up and exhaust your threads.

## A worker-pool pattern (illustrative)

Bounded concurrency keeps you under the rate limit. Not run here because it would call a real API.

```python
from concurrent.futures import ThreadPoolExecutor

def summarise(doc):
    return client.messages.create(model=MODEL, max_tokens=200, timeout=30,
                                  messages=[{"role": "user", "content": f"Summarise: {doc}"}])

with ThreadPoolExecutor(max_workers=4) as pool:      # at most 4 requests in flight
    results = list(pool.map(summarise, documents))
```

## Read the rate-limit headers

Remaining-quota headers let you slow down before you hit a 429.

**Quiz:** Why cap concurrency instead of sending all requests at once?

- [ ] To make tokens free
- [ ] Because HTTP forbids it
- [x] To stay under rate limits and avoid a wall of 429 errors
- [ ] To avoid keys

*Answer:* To stay under rate limits and avoid a wall of 429 errors. Bounded parallelism respects quotas and smooths load.
