---
title: "Rate limits, errors, retries and backoff"
description: "Why reliability engineering matters for AI APIs AI APIs fail in ordinary and unusual ways: rate limits when traffic spikes, temporary overloads at the…"
url: https://optimizeall.com/learn/ai-platform-apis-integration/rate-limits-retries-and-backoff
updated: 2026-10-05
---

Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API · Cost, scale and reliability · lesson 11 of 19 · 15 min

# Rate limits, errors, retries and backoff

## Why reliability engineering matters for AI APIs

AI APIs fail in ordinary and unusual ways: rate limits when traffic spikes, temporary overloads at the provider, network timeouts on long generations, validation errors from malformed requests, and occasionally model refusals. A production integration expects all of them and responds correctly to each.

## Know your error classes

| Class | Examples | Retry? |
|---|---|---|
| Rate limit | HTTP 429 (all providers); Gemini `RESOURCE_EXHAUSTED` | Yes, after waiting (respect `retry-after`) |
| Provider overload / server error | 500, 502, 503, 504; Anthropic 529 `overloaded_error` | Yes, with backoff |
| Timeout / connection error | Network resets, client timeouts | Yes, with backoff (make the operation idempotent) |
| Bad request | 400 (invalid parameters, context too long), 413 (too large) | No: fix the request |
| Auth / permission | 401, 403 | No: fix credentials or access |
| Not found | 404 (unknown model or endpoint) | No: fix configuration |
| Content outcome | Model refusal (e.g., `stop_reason: "refusal"` on some Claude models) or safety block | No blind retry; handle as a product case |

## Rate limits: how they work

Providers limit requests per minute, input tokens per minute and output tokens per minute, per organization or project and per model, and raise limits as your usage and spending tier grow. Response headers tell you where you stand (for example Anthropic's `anthropic-ratelimit-*` headers and OpenAI's `x-ratelimit-*` headers), and 429 responses often include `retry-after`. Check your current limits in each console; don't hard-code them.

## Backoff done right

- **Exponential backoff with jitter**: wait base × 2^attempt plus randomness, capped (e.g., 1 s, 2 s, 4 s… up to 30–60 s). Jitter prevents many clients retrying in lockstep.
- **Respect `retry-after`** when present.
- **Cap attempts and total time**, then fail gracefully (fallback model, queue for later, or a clear message).
- **Idempotency**: only retry operations that are safe to repeat; for side-effecting workflows, use idempotency keys (see the agents course).

The official SDKs already retry some errors: the Anthropic and OpenAI Python SDKs retry connection errors, 408, 409, 429 and 5xx responses twice by default with backoff, configurable via `max_retries`; timeouts are configurable too (units differ by SDK: seconds in Python, milliseconds in TypeScript for Anthropic). Don't stack your own aggressive retry loop on top without accounting for SDK retries, or a single failure can multiply into many calls.

## Hands-on: a retry wrapper with typed errors (Anthropic Python)

```python
import random, time, logging
import anthropic

log = logging.getLogger("ai")
client = anthropic.Anthropic(max_retries=0, timeout=60.0)   # we handle retries ourselves here

RETRYABLE = (anthropic.RateLimitError, anthropic.InternalServerError,     # InternalServerError covers 5xx incl. 529
             anthropic.APIConnectionError, anthropic.APITimeoutError)

def call_with_backoff(max_attempts: int = 5, base: float = 1.0, cap: float = 30.0, **kwargs):
    for attempt in range(max_attempts):
        try:
            return client.messages.create(**kwargs)
        except anthropic.BadRequestError as e:          # 400: our bug, don't retry
            log.error("bad request: %s", e.message)
            raise
        except anthropic.AuthenticationError:            # 401: config problem
            raise
        except RETRYABLE as e:
            retry_after = None
            if isinstance(e, anthropic.APIStatusError):
                retry_after = e.response.headers.get("retry-after")
            wait = float(retry_after) if retry_after else min(cap, base * 2 ** attempt) * random.uniform(0.5, 1.5)
            log.warning("retryable %s (attempt %d), sleeping %.1fs", type(e).__name__, attempt + 1, wait)
            time.sleep(wait)
        except anthropic.APIStatusError:                 # other 4xx (403, 404, 413...): not retryable
            raise
    raise RuntimeError("AI provider unavailable after retries")
```

The OpenAI Python SDK exposes parallel classes (`openai.RateLimitError`, `openai.APIConnectionError`, `openai.APITimeoutError`, `openai.APIStatusError`); the Google Gen AI SDK raises API errors carrying HTTP status codes. Keep one wrapper per provider inside your adapter layer (module 5).

## Protecting yourself upstream

- **Client-side rate limiting**: a token bucket in front of the provider smooths bursts (especially for batch-like loops and parallel fan-out).
- **Queues**: put non-interactive work on a queue with controlled concurrency.
- **Circuit breaker**: if a provider fails repeatedly, stop sending for a short period and switch to a fallback instead of hammering it.
- **Timeouts**: set sensible client timeouts; use streaming for long outputs.
- **Graceful degradation**: cached answers, a smaller model, or "we're busy, we'll email you" beats an error page.

## Worked example: Black Friday at a Dubai fashion retailer

Product-question traffic spiked 8× (illustrative). Before: 429 storms, SDK retries plus a custom retry loop multiplied calls, and users saw errors. After: a token-bucket limiter per model, SDK retries left at default with the custom loop removed, a circuit breaker that shifted overflow traffic to a second provider's model, and a queue for non-urgent tasks. Error rates stayed low and conversion held.

## Pitfalls

- Retrying 400s (you'll just repeat the bug).
- Retry storms without jitter.
- Double retries (SDK + yours) multiplying load.
- Treating refusals as transient errors.

## Measuring success

Error rate by class, retry counts per request, p95 latency including retries, 429 rate versus traffic, and circuit-breaker activations.

## Video lecture: Rate limits, errors, retries and backoff

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

1. Rate limits and retries
2. Why it matters
3. Classify errors
4. Content outcomes
5. How rate limits work
6. Backoff done right
7. Simple example: the crashing loop
8. Circuit breaker, simply
9. Protect upstream
10. Business example: Black Friday in Dubai
11. Common mistakes
12. Test failure on purpose
13. Deeper: Black Friday (illustrative)
14. Watch me do it: call_with_backoff()
15. Recap + try this now

## Lecture transcript

### Rate limits and retries

Everything works in testing. Then launch day arrives, traffic spikes, and your logs fill with four twenty nines, timeouts and overload errors. Users see error pages, and your retry code makes it worse by hammering the provider. This lesson shows you how to handle rate limits and errors like a professional, so the busiest day of your year is also a calm one.

### Why it matters

Why does this matter? Because AI APIs fail more often and more variously than most APIs you've used. Traffic spikes hit rate limits, providers occasionally get overloaded, long generations time out, and some requests are simply invalid. Think of a busy restaurant kitchen at peak hour. A good kitchen doesn't panic when orders pile up. It has a ticket rail, a clear rule for what to cook first, and a backup plan when the grill goes down.

### Classify errors

First, classify errors. Rate limits, the four twenty nine status, mean slow down and try again later, respecting any retry after header. Provider overloads and server errors, like five hundred, five oh three, or Anthropic's five twenty nine overloaded error, are usually temporary, so retry with backoff. Timeouts and connection errors can be retried if the operation is safe to repeat. But bad requests, the four hundreds, authentication errors and not found errors should never be retried. They mean fix your request or configuration.

### Content outcomes

There's one more class that's easy to mishandle: content outcomes, like a model refusal or a safety block. On some newer Claude models, a refusal comes back as a normal response with a refusal stop reason. Retrying blindly usually gets the same result and wastes money. Treat it as a product case: show a helpful message, route to a human, or use a documented fallback approach.

### How rate limits work

How do rate limits work? Providers limit requests per minute and tokens per minute, per organization or project and per model, and raise limits as your usage tier grows. Response headers tell you how much room you have left, and four twenty nine responses often tell you how long to wait. Check your current limits in each provider's console rather than hard coding numbers from a blog post.

### Backoff done right

Now backoff done right. Wait a base amount, doubling with each attempt, with a cap, and add randomness called jitter so thousands of clients don't retry at the same instant. Respect retry after when the provider sends it. Cap the number of attempts and the total time, then fail gracefully. And here's a subtle trap: the official Anthropic and OpenAI SDKs already retry some errors, twice by default. If you wrap them in your own aggressive retry loop, one failure can multiply into many calls.

### Simple example: the crashing loop

A simple example. A freelancer's script summarizes two hundred articles in a loop. Halfway through, it hits a rate limit and crashes. The fix takes ten minutes: let the SDK retry with its defaults, add a short pause between requests, and catch the rate limit error to wait for the retry after period before continuing. The script now finishes every time, just a little slower during busy periods.

### Circuit breaker, simply

Let's look at a circuit breaker more closely, because the name sounds more complex than it is. It's a small piece of code that counts recent failures for a provider. If failures pass a threshold, say five in thirty seconds, the breaker opens and your code stops calling that provider for a short cool down, sending traffic to a fallback instead. After the cool down, it lets a few test requests through. If they succeed, it closes and normal traffic resumes. Like the breaker in your home's fuse box, it protects the whole system from one failing part.

### Protect upstream

Beyond retries, protect yourself upstream. Put a client side rate limiter, like a token bucket, in front of the provider to smooth bursts. Put non interactive work on a queue with controlled concurrency. Add a circuit breaker: if a provider fails repeatedly, stop sending for a short time and switch to a fallback, instead of hammering it. Set sensible timeouts and stream long outputs. And design graceful degradation, like cached answers, a smaller model, or a we'll email you message.

### Business example: Black Friday in Dubai

Now a realistic business example, with illustrative numbers. On Black Friday, a Dubai fashion retailer's product question traffic jumped about eight times. Before the fix, rate limit storms hit, the SDK's retries plus a custom retry loop multiplied calls, and shoppers saw errors. After the fix: a token bucket limiter per model, the custom retry loop removed, a circuit breaker that shifted overflow to a second provider, and a queue for non urgent tasks. Error rates stayed low and conversion held.

### Common mistakes

Common mistakes. Retrying four hundreds, which just repeats the bug. Retry storms without jitter. Double retries from the SDK plus your own loop. Treating refusals as temporary errors. And no graceful degradation, so the only failure mode is an error page.

### Test failure on purpose

How do you know your error handling works? Test it on purpose. In staging, point your client at a mock server that returns four twenty nines, five hundreds and slow responses on a schedule, or use a load testing tool to exceed your own rate limits. Watch what users would see, how many total calls are made per request, and whether the circuit breaker and fallback kick in. It's far better to discover a retry storm in a controlled test than on your biggest sales day.

### Deeper: Black Friday (illustrative)

Let's deepen the Dubai fashion retailer's Black Friday with illustrative numbers. Traffic reached roughly eight times normal for six hours. In the previous year, 429 errors peaked at a painful share of requests, and duplicate retries from the SDK plus a custom loop multiplied provider calls. This year, the token bucket smoothed bursts, the circuit breaker opened twice for a few minutes, shifting overflow to the second provider, and non urgent tasks like review summaries waited in a queue until the evening. Customer facing errors stayed low throughout, and the engineering team, for once, watched the dashboard instead of firefighting.

### Watch me do it: call_with_backoff()

Watch me do it. Let's walk through the retry wrapper. I create the Anthropic client with SDK retries turned off, because this wrapper owns retries, and a sixty second timeout. Retryable lists rate limit errors, internal server errors, which cover five hundreds including overloaded responses, connection errors and timeouts. The loop tries up to five times. A bad request error is our bug, so I log it and raise immediately. An authentication error means configuration, so I raise. For retryable errors, I read the retry after header when the error carries a response, and if there isn't one, compute exponential backoff capped at thirty seconds, multiplied by a random factor between a half and one and a half. That's the jitter. I log and sleep, then try again. Any other status error, like a four oh three or four oh four, raises. After five attempts, I raise a clear provider unavailable error, which the router turns into a fallback.

### Recap + try this now

Quick recap. Classify errors and only retry what's retryable. Use exponential backoff with jitter, respect retry after, and don't stack retries on top of the SDK's. Protect upstream with limiters, queues and circuit breakers, and degrade gracefully. Try this now: wrap your provider calls with typed error handling and backoff, add a simple token bucket limiter, and run a small load test that triggers rate limits to confirm your app stays calm.

## Key takeaways

- Classify errors: retry rate limits, overloads and timeouts; fix bad requests, auth and not-found errors.
- Use exponential backoff with jitter, respect retry-after, and cap attempts.
- Official SDKs already retry some errors; avoid stacking retry loops that multiply calls.
- Protect upstream with client-side rate limiting, queues, circuit breakers and timeouts.
- Degrade gracefully with fallbacks, smaller models or deferred processing.

## Try it

Wrap your provider calls with typed error handling and backoff, add a simple token-bucket limiter, and run a load test that triggers 429s to confirm graceful behavior.

- [Previous: Prompt caching across providers](https://optimizeall.com/learn/ai-platform-apis-integration/prompt-caching-across-providers)
- [Next: Cost tracking, attribution and budgets](https://optimizeall.com/learn/ai-platform-apis-integration/cost-tracking-and-budgets)
- [All lessons of Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API](https://optimizeall.com/learn/ai-platform-apis-integration)
