Skip to content

Integrating AI Platforms: Claude, OpenAI, Gemini and Open Models via API · Cost, scale and reliability · lesson 9 of 19 · 14 min

Batch APIs: bulk processing at a discount

When batch beats real time

Many AI jobs don't need an answer in seconds: enriching 50,000 CRM records overnight, classifying a month of support tickets, generating product descriptions for a new catalog, running an evaluation suite, embedding a document archive. Batch APIs accept many requests at once, process them asynchronously (typically within hours; many batches finish much sooner) and charge less. At the time of writing, Anthropic's Message Batches and OpenAI's Batch API are both priced at 50% of standard rates, and Gemini offers a discounted batch mode; check current pricing pages.

How each provider's batch works

| | Claude Message Batches | OpenAI Batch API | Gemini batch | |---|---|---|---| | Submit | client.messages.batches.create(requests=[{"custom_id", "params"}]) | Upload a JSONL file (purpose="batch"), then client.batches.create(input_file_id=..., endpoint="/v1/responses", completion_window="24h") | client.batches.create(model=..., src=...) with inline requests (Developer API) or a Cloud Storage / BigQuery source on the enterprise platform | | Track | batches.retrieve(id).processing_status until ended | batches.retrieve(id).status until completed | batches.get(name=...) state | | Results | Stream batches.results(id); each has custom_id and result.type (succeeded, errored, canceled, expired) | Download the output file (and error file); lines keyed by custom_id | Inline responses or an output location | | Limits | Up to 100,000 requests or 256 MB per batch; most finish within an hour, maximum 24 h (at time of writing) | Endpoint-specific limits (e.g., embeddings batches capped by total inputs) | See docs | | Availability | Claude API and Claude Platform on AWS (not every cloud platform) | OpenAI API | Developer API and enterprise platform (different sources) |

Results can arrive in any order: always key by your own custom_id.

Hands-on: Claude batch for product descriptions

import os, time, json
import anthropic
from anthropic.types.message_create_params import MessageCreateParamsNonStreaming
from anthropic.types.messages.batch_create_params import Request

client = anthropic.Anthropic()
MODEL = os.environ.get("CLAUDE_MODEL", "claude-sonnet-5")
products = [{"sku": "LHR-001", "name": "Handwoven khes throw", "facts": "cotton, 150x200cm, Punjab artisans"},
            {"sku": "LHR-002", "name": "Brass tea set", "facts": "4 cups, hand-engraved, food-safe lining"}]

batch = client.messages.batches.create(requests=[
    Request(custom_id=p["sku"], params=MessageCreateParamsNonStreaming(
        model=MODEL, max_tokens=400,
        system="Write a 60-word product description in warm, plain English. No unverifiable claims.",
        messages=[{"role": "user", "content": f"{p['name']}: {p['facts']}"}]))
    for p in products])

while (b := client.messages.batches.retrieve(batch.id)).processing_status != "ended":
    time.sleep(60)                                   # poll politely; use a scheduler in production

out = {}
for r in client.messages.batches.results(batch.id):
    if r.result.type == "succeeded":
        out[r.custom_id] = "".join(x.text for x in r.result.message.content if x.type == "text")
    else:
        out[r.custom_id] = {"error": r.result.type}   # re-queue errored/expired items
print(json.dumps(out, indent=2, ensure_ascii=False))

Hands-on: OpenAI batch file (JSONL)

import json
from openai import OpenAI
client = OpenAI()
with open("batch.jsonl", "w") as f:
    for p in products:
        f.write(json.dumps({"custom_id": p["sku"], "method": "POST", "url": "/v1/responses",
                            "body": {"model": "gpt-5.5", "instructions": "Write a 60-word product description.",
                                     "input": f"{p['name']}: {p['facts']}"}}) + "\n")
up = client.files.create(file=open("batch.jsonl", "rb"), purpose="batch")
job = client.batches.create(input_file_id=up.id, endpoint="/v1/responses", completion_window="24h")
print(job.id, job.status)   # later: client.batches.retrieve(job.id); download output_file_id content

Designing batch pipelines

  • Idempotent submission: record which items were submitted in which batch; never resubmit blindly.
  • Partial failure handling: collect errored and expired items and resubmit them in a follow-up batch (perhaps with a different model).
  • Validation: run the same structured-output validation as real-time paths.
  • Scheduling: submit in the evening, process results in the morning; notify owners when done.
  • Combine with caching: shared system prompts in batch requests can also benefit from prompt caching where supported, stacking discounts.
  • Don't batch user-facing work: anything a person is waiting on belongs in real-time APIs.

Worked example: a Karachi marketplace catalog refresh

A marketplace adds 30,000 new products per month from sellers' spreadsheets. Previously they generated descriptions in real time, hitting rate limits during business hours. Now a nightly batch handles new SKUs with structured outputs (title, description, bullet features, category). Morning validation flags ~2% for human review (illustrative), and costs roughly halve versus real-time calls at the same model.

Cost math you can show a manager

A quick way to decide: estimate monthly items × average input and output tokens × price per token, for real-time and batch. Add the cost of the engineering needed to schedule, validate and retry, and subtract the cost of rate-limit congestion you avoid during business hours (other features slowing down, retries). For recurring bulk jobs above a few thousand items a month, batch usually wins clearly; for small or urgent jobs, the operational overhead may not be worth it. Write the numbers down with the date and pricing page you used, because prices change.

Pitfalls

  • Relying on result order.
  • No plan for expired or errored items.
  • Treating the batch window as a guarantee for time-critical work.
  • Forgetting that batch availability differs by platform.

Measuring success

Cost per item versus real-time, completion time distribution, error/expiry rate, and validation failure rate.

Video lecture: Batch APIs: bulk processing at a discount

Lecture coming soon · 15 chapters · about 9 minutes. Read the full transcript below.

  1. Batch APIs
  2. Why it matters
  3. How batch works
  4. Provider flows
  5. Key by custom_id
  6. Simple example: 80 cake descriptions
  7. Business example: Karachi marketplace
  8. Is batch worth it?
  9. Pipeline design
  10. Availability varies
  11. Common mistakes
  12. Make batches observable
  13. Deeper: the marketplace batch (illustrative)
  14. Watch me do it: Claude batch
  15. Recap + try this now

Lecture transcript

Batch APIs

Imagine you need product descriptions for thirty thousand new items, or you want to classify a month of support tickets. You could fire thousands of real time requests, hit rate limits all afternoon, and pay full price. Or you could hand the whole pile to a batch API, go home, and collect the results in the morning at a discount. In this lesson you'll learn when batch is the right call, and how to run it with Claude, OpenAI and Gemini.

Why it matters

Why does this matter? Because a surprising share of AI work isn't interactive at all. Enrichment, classification, content generation for catalogs, evaluation runs and embedding archives can all wait a few hours. Think of posting letters versus making phone calls. A phone call is immediate but you pay for your time. A letter takes longer, but you can send a thousand at once, cheaply. Batch is the postal service of AI APIs.

How batch works

How does it work? You submit many requests at once, each with your own id. The provider processes them in the background, usually within hours and often much sooner, and you collect the results when they're done. At the time of writing, Anthropic's message batches and OpenAI's batch API are both priced at half of standard rates, and Gemini has a discounted batch mode too. Check the current pricing pages before you plan budgets.

Provider flows

With Claude, you call message batches create with a list of requests, each with a custom id and the same parameters you'd send to messages create. You poll until the processing status is ended, then stream the results. Each result has your custom id and a result type: succeeded, errored, canceled or expired. With OpenAI, you write a JSON lines file where each line is a request with a custom id, upload it with the batch purpose, then create a batch pointing to that file and an endpoint like the Responses API. Gemini accepts inline requests on its developer API and cloud storage or BigQuery sources on its enterprise platform.

Key by custom_id

The golden rule: results can come back in any order. Never match results by position. Always use your own custom id, like the product SKU or ticket number. That one habit prevents the nightmare scenario where product A gets product B's description, and nobody notices until customers complain.

Simple example: 80 cake descriptions

A simple example. A small bakery chain wants a short, friendly description for each of its eighty cakes. Instead of generating them one by one during the day, the owner submits one batch in the evening with each cake's name and ingredients, keyed by its code. In the morning, eighty descriptions are waiting. Two came back errored because the ingredient list was empty, so she fixes those two and resubmits. The whole job cost about half what real time would have.

Business example: Karachi marketplace

Now a realistic business example, with illustrative numbers. A Karachi marketplace adds about thirty thousand products a month from sellers' spreadsheets. They used to generate descriptions in real time and hit rate limits during business hours, slowing everything else down. Now a nightly batch handles new SKUs with structured outputs: title, description, feature bullets and category. Morning validation flags a small share for human review, and costs roughly halve versus real time calls on the same model.

Is batch worth it?

How do you decide whether batch is worth it? Do the simple math. Monthly items, times average input and output tokens, times the price per token, once for real time and once for batch. Then add the engineering effort for scheduling, validation and retries, and subtract the cost of rate limit congestion you avoid during the day, when other features slow down. For recurring jobs above a few thousand items a month, batch usually wins clearly. For small or urgent jobs, the overhead may not pay off. Write the numbers down with the date and the pricing page you used, because prices change.

Pipeline design

Design your pipeline for reality. Record which items went into which batch, so you never resubmit blindly. Collect errored and expired items and resubmit them in a follow up batch, perhaps with a different model. Validate outputs exactly as you would in real time. Schedule submissions in the evening and notify owners when results land. And where supported, stack prompt caching with batch for shared system prompts.

Availability varies

Also check availability. Batch features differ by platform: Claude's message batches are available on the Claude API and Claude Platform on AWS, but not on every cloud platform version of Claude. OpenAI's batch endpoint supports a specific list of API endpoints. Gemini's batch sources differ between its developer API and enterprise platform. Always confirm in the docs for the exact platform you deploy on.

Common mistakes

Common mistakes. Relying on result order. Having no plan for expired or errored items. Treating the batch window as a guarantee for time critical work. And batching anything a person is actively waiting for. If a user is staring at a screen, use real time and streaming instead.

Make batches observable

One more operational tip: make batches observable. Log when each batch was submitted, how many items it contained, when it ended, and how many succeeded, errored or expired. Alert if a batch hasn't ended well within the provider's window, or if the error rate jumps. And keep the raw inputs for a few days, so you can resubmit quickly if something goes wrong. Batch jobs run while you sleep, so you want the morning to start with a clear summary, not a mystery.

Deeper: the marketplace batch (illustrative)

Let's deepen the Karachi marketplace example with illustrative numbers. About thirty thousand new products a month, arriving in sellers' spreadsheets. The nightly batch runs on new SKUs only, using structured outputs for title, description, feature bullets and category. Most requests succeed; a small number error on empty or malformed spec sheets and go back to sellers with a request to complete their data. Morning validation flags a couple of percent for human review, mostly items with claims like waterproof or organic that need evidence. Cost per description roughly halved versus real time, and daytime rate limit errors disappeared, so the customer facing assistant got faster too.

Watch me do it: Claude batch

Watch me do it. Let's walk through the Claude batch script. The products list has two items with a SKU, a name and facts. I create a batch with one request per product. Each request has a custom id, the SKU, and params exactly like a normal messages create: model, max tokens, a system prompt asking for sixty warm words without unverifiable claims, and one user message with the name and facts. The loop retrieves the batch every sixty seconds until the processing status is ended; in production a scheduler does this. Then I stream the results. For each one, if the result type is succeeded, I join the text blocks and store them under the custom id. Otherwise, I store the error type, so I can resubmit those items later. The OpenAI version writes one JSON line per request with a custom id, method, URL and body, uploads the file with the batch purpose, and creates the batch against the Responses endpoint.

Recap + try this now

Quick recap. Batch APIs process big, non urgent jobs asynchronously at a discount. Claude takes request lists, OpenAI takes JSON lines files, and Gemini takes inline or cloud sources. Key by custom id, handle failures, validate outputs, and keep user facing work real time. Try this now: pick one non urgent workload with at least two hundred items, move it to a batch API with custom id tracking and a retry path, and compare the cost and turnaround with your real time version.

Key takeaways

  • Batch APIs process large volumes asynchronously at a discount; use them for non-urgent work.
  • Claude batches take request lists; OpenAI batches take uploaded JSONL files; Gemini batches take inline or cloud-storage sources.
  • Results can arrive in any order: key everything by your own custom_id.
  • Plan for errored and expired items with follow-up batches and validation.
  • Never batch work a user is waiting on; check platform availability.

Try it

Move one non-urgent workload (at least 200 items) to a batch API, add custom_id tracking and a retry path for failures, and compare cost and turnaround with your real-time version.