Mastering ChatGPT (OpenAI)The OpenAI API and business integrations · Lesson 16 of 19
The OpenAI API: Responses API fundamentals with working code
Video lecture
The OpenAI API: Responses API fundamentals with working code
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 The OpenAI API
ChatGPT is brilliant when a person is typing. But when three thousand products need descriptions, or every support ticket needs tagging, you want software doing the work. That is the OpenAI API. In this lecture you will learn the building blocks, make a first call, get validated JSON back, control cost, and learn what to put in place before launch.
0:26 Why the API matters
Why learn the API? Because ChatGPT works at the speed of a person, and some jobs need the speed of a process. Three thousand product descriptions. Every support ticket tagged within a minute. Every call note summarised into the CRM. The API puts the same families of models inside your own systems. Think of ChatGPT as hiring a brilliant consultant by the hour, and the API as building their skill into your assembly line.
0:58 App vs API
ChatGPT is an app for people. The OpenAI API lets your software use the same families of models. You sign up on the OpenAI Platform, create a project and an API key, add billing, and pay per use, based on tokens in and out, priced per model. Importantly, ChatGPT subscriptions and API billing are completely separate.
1:22 Building blocks
The primary interface is the Responses API. It generates text and uses tools, built in ones like web search, file search, code interpreter and remote MCP servers, plus your own functions. It handles structured outputs, image inputs and conversation state. You will still see the older Chat Completions API in tutorials, and it remains supported. Models have exact IDs, so check the models page for current names and prices. Instructions carry your system level rules, and reasoning models accept an effort setting.
1:58 First call
Here is the pattern from the lesson. Install the official Python library. Put your key in an environment variable, never in code. Create a client, which reads the key automatically. Call responses dot create with a model, instructions, for example a dates brand that forbids health claims and marks missing facts as T B C, and the input. Wrap it in error handling for rate limits, API errors and connection problems. Then read output text, and log the input and output tokens.
2:34 Structured outputs
When code needs reliable data, use structured outputs. Define a Pydantic model, say a review tag with sentiment, topic, needs reply and a summary, where sentiment and topic must come from fixed lists. Call responses dot parse with that model as the text format. You get back a validated object, not a paragraph you have to pick apart.
2:59 Cost control
Now cost. Use the smallest model that passes your quality bar. Test reasoning effort at low, medium and high on your evaluation set, and keep the lowest level that holds quality. Keep long, stable prompt prefixes identical so caching can help where supported. And send non urgent bulk work through the Batch API, which is cheaper. Set budget alerts and project spend limits in the dashboard too.
3:28 Before launch
Before launch, five things. An evaluation set of thirty to a hundred real examples with expected outputs, rerun on every change. Guardrails in instructions, schemas and code. A human hand off for sensitive cases. Key hygiene, meaning server side only, per project keys, least privilege and rotation if a key is exposed. And monitoring of cost, latency, errors and quality.
3:54 Simple example
A simple example. Make one call that writes a sixty word description for one product, and read the usage, the input and output token counts. Now multiply by three thousand products and by the model's current price from the pricing page. You have a cost estimate before you process anything real. If it is too high, try a smaller model or lower reasoning effort on your evaluation examples. This habit alone prevents most surprise API bills.
4:27 Worked example: 3,000 bilingual descriptions
A Gulf retailer needs Arabic and English descriptions for three thousand products. The team tests the prompt on fifty products with a native speaker reviewing, then adds a schema with English and Arabic fields and a missing information list. The full catalogue runs through the Batch API overnight. Items with missing information go to merchandisers, and everything else is spot checked, five percent, before publishing. The team logged token usage on the first fifty products and multiplied by the catalogue size to estimate the batch cost before running it. The estimate was within their budget, but they still switched the Arabic review step to a sample of ten percent after the first week, once the native reviewer's error rate had dropped close to zero. Measure first, then decide where people add the most value.
5:25 Pitfalls
Four pitfalls. Calling the API from browser code with a visible key. Assuming your ChatGPT subscription covers API usage. Launching without an evaluation set. And ignoring incomplete or failed responses. Handle them, log them, and route them to a person.
5:42 Try this now
Try this now. On the OpenAI Platform, create a project and an API key, add a spend limit, and store the key in an environment variable. Run the first call example on a non sensitive product. Then take ten real, anonymised reviews and run them through the Pydantic structured output example. Check every classification by hand, and where you disagree, add one clear rule or example to the instructions and run the ten again. You have just built the start of an evaluation set.
6:19 Common mistakes
Here are the mistakes that cause most API pain. Putting a key in browser code, where anyone can copy it. Forgetting to set a project spend limit, so a bug in a loop becomes a big bill. Launching without an evaluation set, so nobody knows whether a prompt change made things better or worse. And ignoring failed or incomplete responses, which then show up as blank descriptions on your website. Each one takes minutes to prevent.
6:52 Watch me do it, part 1
Let me run the dates brand example. On the OpenAI Platform I create a project, generate a key, which I keep masked, and set a monthly spend limit before anything else. In the terminal I export the key and the model name. In the editor, the client reads the key automatically. The instructions say British English, no health claims, T B C for missing information. responses dot create sends the product, and three except blocks catch rate limits, API errors and connection problems. I run it. A description for the Medjool gift box prints, the origin marked T B C because I left it blank, and below it the token counts.
7:40 Watch me do it, part 2
Now structured output. The ReviewTag class uses literal types, so sentiment and topic can only be values from fixed lists. I call responses dot parse with that class as the text format, and loop over ten anonymised reviews. Each result is a validated object, not text. I compare with my own labels. One courier complaint is tagged service, so I add a rule to the instructions, and the rerun matches all ten. Then a cost test. I set reasoning effort to low and rerun. Still ten out of ten, with fewer output tokens. For this task, low effort is the right setting.
8:24 Recap and next step
Recap. Use the Responses API with instructions and a current model ID, keep keys in environment variables, parse structured outputs into validated objects, tune model size and reasoning effort for cost, and launch only with evaluation, guardrails, hand off and monitoring. Your next step: run the first call example with your own key, then classify ten real, anonymised reviews with the Pydantic schema and check each one by hand.
ChatGPT vs the OpenAI API
ChatGPT is an app for people. The OpenAI API lets your software use OpenAI's models: classify every support ticket, generate product descriptions from a catalogue, summarise call notes into your CRM, or power a customer-facing assistant. You create an account on the OpenAI Platform, create a project and an API key, add billing, and pay per use (tokens in and out, priced per model). ChatGPT subscriptions and API billing are separate.
The core building blocks
- Responses API (
client.responses.create): OpenAI's primary API for generating text and using tools. It supports built-in tools (such as web search, file search, code interpreter and remote MCP servers), function calling, structured outputs, images as input, and conversation state (previous_response_idor the Conversations API). - Chat Completions API: the previous standard, still supported; you will see it in older tutorials.
- Models: exact IDs such as
gpt-5.5or newer GPT-6-generation models. Check the models page for current IDs, capabilities and prices; use the smallest model that meets your quality bar. - Instructions: system-level guidance (the API equivalent of custom instructions).
- Reasoning effort: reasoning models accept an effort setting (for example low, medium, high) that trades depth for speed and cost.
- Structured outputs: constrain output to a JSON schema.
- Batch API and flex/priority options: process large jobs asynchronously at lower cost, or pay for priority. Check the pricing page for current discounts.
Hands-on: first call in Python
pip install openai
export OPENAI_API_KEY="sk-..." # from the OpenAI Platform; never commit it
export OPENAI_MODEL="gpt-5.5" # check the models page for current IDsimport os
import openai
from openai import OpenAI
client = OpenAI() # reads OPENAI_API_KEY from the environment
MODEL = os.environ.get("OPENAI_MODEL", "gpt-5.5")
INSTRUCTIONS = (
"You write product descriptions for Al Waha Dates (premium dates, UAE/UK). "
"British English. No health claims. Never invent origins, awards or "
"ingredients; write [TBC] if information is missing."
)
def describe(product: dict) -> str:
try:
resp = client.responses.create(
model=MODEL,
instructions=INSTRUCTIONS,
input=f"Write a 60-word description for this product:\n{product}",
)
except openai.RateLimitError:
raise RuntimeError("Rate limited - slow down or retry later")
except openai.APIStatusError as e:
raise RuntimeError(f"OpenAI API error {e.status_code}")
except openai.APIConnectionError:
raise RuntimeError("Network problem reaching the OpenAI API")
print("tokens in/out:", resp.usage.input_tokens, resp.usage.output_tokens)
return resp.output_text
if __name__ == "__main__":
print(describe({"name": "Medjool Gift Box", "weight": "500g", "origin": "[TBC]"}))Structured outputs with Pydantic
from typing import Literal
from pydantic import BaseModel
class ReviewTag(BaseModel):
sentiment: Literal["positive", "neutral", "negative"]
topic: Literal["delivery", "product", "price", "service", "other"]
needs_reply: bool
summary: str
resp = client.responses.parse(
model=MODEL,
input=f"Classify this customer review:\n{review_text}",
text_format=ReviewTag,
)
tag = resp.output_parsed # a validated ReviewTag instanceReasoning effort and cost control
For analysis-heavy tasks, set reasoning effort explicitly and measure:
resp = client.responses.create(
model=MODEL,
reasoning={"effort": "low"}, # try low/medium/high on your evaluation set
input="Summarise these call notes into 3 action items: ...",
)Cost levers: smaller models for simple tasks, lower effort where quality holds, prompt caching of long stable prefixes (automatic for repeated prefixes on supported models; check the docs), and the Batch API for non-urgent bulk work.
Before you launch
- Evaluation set: 30 to 100 real examples with expected outputs; rerun on every prompt or model change.
- Guardrails: instructions, schemas, and code checks for banned claims or personal data.
- Human hand-off for sensitive cases.
- Key hygiene: server-side only, per-project keys, least-privilege permissions, rotation on exposure.
- Monitoring: usage, cost, latency, errors and quality; set budget alerts and project spend limits in the dashboard.
Worked example: catalogue descriptions for a Gulf retailer
A retailer with 3,000 products needs Arabic and English descriptions. The team tests the prompt on 50 products with a native-speaker reviewer, adds a schema with en and ar fields and a missing_info list, then runs the full catalogue through the Batch API overnight. Items with missing_info go to merchandisers; everything else is spot-checked (5%) before publishing.
Keeping conversation state
For multi-turn assistants, either pass the previous turns back in input, or pass previous_response_id=resp.id so the API continues from the prior response. Store only what you need, and remember that longer histories cost more tokens on every turn.
Pitfalls
- Calling the API from browser code with a visible key.
- Confusing ChatGPT plans with API billing.
- No evaluation set; "vibes-based" launches.
- Ignoring incomplete responses (check status and handle errors).
How to measure success
Your evaluation accuracy meets an agreed bar, cost per item is known, failures are handled gracefully, and keys are never exposed.
Key takeaways
- The OpenAI API lets software use OpenAI models; it is billed separately from ChatGPT plans, per token and model.
- The Responses API is the primary interface (tools, function calling, structured outputs, state); Chat Completions remains supported.
- Use responses.parse with a Pydantic model for validated JSON; tune model size and reasoning effort; use Batch for bulk work.
- Launch with an evaluation set, guardrails, human hand-off, server-side keys, spend limits and monitoring.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Run the first-call example with your own key on a non-sensitive task, then classify 10 anonymised reviews with the Pydantic schema, checking each by hand and noting any misclassification.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.