Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via APICost, scale and reliability · Lesson 11 of 19

Rate limits, errors, retries and backoff

Article · 15 min · 9 min lecture

Video lecture

Rate limits, errors, retries and backoff

15 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 15

Rate limits and retries

  • Error classes
  • Backoff done right
  • Protecting upstream
  • Graceful degradation

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Why reliability engineering matters for AI APIs

AI APIs fail in ordinary and unusual ways: rate limits when traffic spikes, temporary overloads at the provider, network timeouts on long generations, validation errors from malformed requests, and occasionally model refusals. A production integration expects all of them and responds correctly to each.

Know your error classes

ClassExamplesRetry?
Rate limitHTTP 429 (all providers); Gemini RESOURCE_EXHAUSTEDYes, after waiting (respect retry-after)
Provider overload / server error500, 502, 503, 504; Anthropic 529 overloaded_errorYes, with backoff
Timeout / connection errorNetwork resets, client timeoutsYes, with backoff (make the operation idempotent)
Bad request400 (invalid parameters, context too long), 413 (too large)No: fix the request
Auth / permission401, 403No: fix credentials or access
Not found404 (unknown model or endpoint)No: fix configuration
Content outcomeModel refusal (e.g., stop_reason: "refusal" on some Claude models) or safety blockNo blind retry; handle as a product case

Rate limits: how they work

Providers limit requests per minute, input tokens per minute and output tokens per minute, per organization or project and per model, and raise limits as your usage and spending tier grow. Response headers tell you where you stand (for example Anthropic's anthropic-ratelimit-* headers and OpenAI's x-ratelimit-* headers), and 429 responses often include retry-after. Check your current limits in each console; don't hard-code them.

Backoff done right

  • Exponential backoff with jitter: wait base × 2^attempt plus randomness, capped (e.g., 1 s, 2 s, 4 s… up to 30–60 s). Jitter prevents many clients retrying in lockstep.
  • Respect retry-after when present.
  • Cap attempts and total time, then fail gracefully (fallback model, queue for later, or a clear message).
  • Idempotency: only retry operations that are safe to repeat; for side-effecting workflows, use idempotency keys (see the agents course).

The official SDKs already retry some errors: the Anthropic and OpenAI Python SDKs retry connection errors, 408, 409, 429 and 5xx responses twice by default with backoff, configurable via max_retries; timeouts are configurable too (units differ by SDK: seconds in Python, milliseconds in TypeScript for Anthropic). Don't stack your own aggressive retry loop on top without accounting for SDK retries, or a single failure can multiply into many calls.

Hands-on: a retry wrapper with typed errors (Anthropic Python)

import random, time, logging
import anthropic

log = logging.getLogger("ai")
client = anthropic.Anthropic(max_retries=0, timeout=60.0)   # we handle retries ourselves here

RETRYABLE = (anthropic.RateLimitError, anthropic.InternalServerError,     # InternalServerError covers 5xx incl. 529
             anthropic.APIConnectionError, anthropic.APITimeoutError)

def call_with_backoff(max_attempts: int = 5, base: float = 1.0, cap: float = 30.0, **kwargs):
    for attempt in range(max_attempts):
        try:
            return client.messages.create(**kwargs)
        except anthropic.BadRequestError as e:          # 400: our bug, don't retry
            log.error("bad request: %s", e.message)
            raise
        except anthropic.AuthenticationError:            # 401: config problem
            raise
        except RETRYABLE as e:
            retry_after = None
            if isinstance(e, anthropic.APIStatusError):
                retry_after = e.response.headers.get("retry-after")
            wait = float(retry_after) if retry_after else min(cap, base * 2 ** attempt) * random.uniform(0.5, 1.5)
            log.warning("retryable %s (attempt %d), sleeping %.1fs", type(e).__name__, attempt + 1, wait)
            time.sleep(wait)
        except anthropic.APIStatusError:                 # other 4xx (403, 404, 413...): not retryable
            raise
    raise RuntimeError("AI provider unavailable after retries")

The OpenAI Python SDK exposes parallel classes (openai.RateLimitError, openai.APIConnectionError, openai.APITimeoutError, openai.APIStatusError); the Google Gen AI SDK raises API errors carrying HTTP status codes. Keep one wrapper per provider inside your adapter layer (module 5).

Protecting yourself upstream

  • Client-side rate limiting: a token bucket in front of the provider smooths bursts (especially for batch-like loops and parallel fan-out).
  • Queues: put non-interactive work on a queue with controlled concurrency.
  • Circuit breaker: if a provider fails repeatedly, stop sending for a short period and switch to a fallback instead of hammering it.
  • Timeouts: set sensible client timeouts; use streaming for long outputs.
  • Graceful degradation: cached answers, a smaller model, or "we're busy, we'll email you" beats an error page.

Worked example: Black Friday at a Dubai fashion retailer

Product-question traffic spiked 8× (illustrative). Before: 429 storms, SDK retries plus a custom retry loop multiplied calls, and users saw errors. After: a token-bucket limiter per model, SDK retries left at default with the custom loop removed, a circuit breaker that shifted overflow traffic to a second provider's model, and a queue for non-urgent tasks. Error rates stayed low and conversion held.

Pitfalls

  • Retrying 400s (you'll just repeat the bug).
  • Retry storms without jitter.
  • Double retries (SDK + yours) multiplying load.
  • Treating refusals as transient errors.

Measuring success

Error rate by class, retry counts per request, p95 latency including retries, 429 rate versus traffic, and circuit-breaker activations.

Key takeaways

  • Classify errors: retry rate limits, overloads and timeouts; fix bad requests, auth and not-found errors.
  • Use exponential backoff with jitter, respect retry-after, and cap attempts.
  • Official SDKs already retry some errors; avoid stacking retry loops that multiply calls.
  • Protect upstream with client-side rate limiting, queues, circuit breakers and timeouts.
  • Degrade gracefully with fallbacks, smaller models or deferred processing.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. Your request fails with HTTP 400 'context too long'. What should your code do?
  2. Why add jitter to exponential backoff?
  3. After adding a custom retry loop, load during outages multiplies. What is a likely cause?

Put it into practice

Wrap your provider calls with typed error handling and backoff, add a simple token-bucket limiter, and run a load test that triggers 429s to confirm graceful behavior.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.