Mastering Claude (Anthropic)Building with the Claude API · Lesson 16 of 20
The Claude API: core concepts and your first working integration
Video lecture
The Claude API: core concepts and your first working integration
The narrated lecture is in production
Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.
Chapters
Transcript of the narration, chapter by chapter.
0:00 The Claude API
The Claude app is perfect when a person is involved in every task. But what about the two thousand reviews that arrive every month, or every new lead from your website, or a whole product catalogue that needs descriptions? That is where the API comes in. In this lecture you will learn the core concepts, see your first working call, get reliable JSON back, and learn what to put in place before you launch.
0:32 Why the API matters
Why learn the API if the app already works? Because the app helps a person, while the API helps a process. When the same job must happen hundreds or thousands of times, tagging reviews, summarising tickets, writing descriptions, a person in a chat window becomes the bottleneck. The API lets your software do it, consistently, around the clock. Think of the app as a skilled freelancer at your desk, and the API as that same skill built into your production line.
1:07 App or API?
Use the app when a person is in the loop. Use the API when software should call Claude automatically. You sign up in the Claude Console, create an API key, and pay per use, billed in tokens, which are roughly word pieces, with different prices for input and output depending on the model. If your company already buys cloud from Amazon, Google or Microsoft, Claude is also available through their platforms.
1:38 Core concepts
The core is simple. One endpoint called the Messages API. You send a model name, a system prompt that sets role and rules, like Project instructions, a list of messages, and a maximum number of output tokens. You get back content blocks and a usage count. Model identifiers are exact strings, like claude sonnet five or claude haiku four five. Always check the models page for the current list, because new models arrive regularly. Current models also support adaptive thinking, with an effort setting that trades depth for speed and cost.
2:18 Your first call
Here is the pattern from the lesson. Install the official Python SDK. Put your API key in an environment variable, never in the code. Create a client, which reads that variable automatically. Call messages dot create with the model, a system prompt, say for a skincare brand that forbids medical claims and marks missing facts as T B C, and a user message with the product data. Wrap it in error handling for rate limits, API errors and network problems, and log the input and output tokens so you know your cost per item.
2:59 Structured outputs
Free text is hard for software to use. Structured outputs fix that. You define a JSON schema, for example sentiment must be positive, neutral or negative, topic must be one of five categories, needs reply is true or false, plus a short summary. Pass that schema in the output configuration, and Claude's response is constrained to match it. Now your code can write straight into a sheet or a CRM without fragile parsing.
3:31 Tools and cost levers
Two more concepts. Tool use lets Claude ask your application to run a function, like look up an order, and your code sends back the result. We build that in the next lesson. And cost levers. Prompt caching makes a long, stable prefix, like a big style guide, much cheaper to reuse. The Message Batches API processes large jobs asynchronously at a discount, currently half price. And the biggest lever of all is choosing the smallest model that passes your quality bar.
4:07 Before launch
Before you launch, five things. An evaluation set of thirty to a hundred real examples with expected outputs, rerun on every change. Guardrails in the prompt, the schema and your code. A human hand off for sensitive cases like complaints, legal threats or health questions. Secrets on the server only, never in browser code or a repository, and rotated immediately if exposed. And monitoring of cost, latency, errors and quality.
4:37 Simple example
A simple example. One call, one job. Send the product facts for a rose hydrating mist, a system prompt saying British spelling and no medical claims, and ask for a sixty word description. The response comes back with the text and a usage count, for instance, around one hundred and eighty input tokens and ninety five output tokens. Multiply by your catalogue size and your model's price, and you have a cost estimate before you process a single real product. That habit, measuring one item first, prevents surprise bills.
5:16 Worked example: 2,000 reviews a month
An electronics seller in Karachi gets about two thousand marketplace reviews a month. A seventy line script sends each review to a fast model with the schema, writes the results to a Google Sheet, and flags the ones needing a reply for the support lead. Before going live, they tested the prompt on one hundred reviews they had labelled by hand. Then they moved the old back catalogue to the Batches API to cut costs. And every negative product review is still read by a human.
5:53 Try this now
Try this now. Create an API key in the Claude Console and store it in an environment variable. Run the first call example from the lesson on a non sensitive product or topic, and note the input and output token counts. Then take ten real reviews, remove any names or order numbers, and run them through the structured output example with the sentiment and topic schema. Check every result by hand. How many did you agree with? Where you disagree, improve the system prompt with one clear rule or example and run the ten again. You have just built your first tiny evaluation set.
6:38 Watch me do it, part 1
Let me run the first call on screen. In the terminal I install the anthropic package and set two environment variables, my API key, which I mask, and the model name. In the editor, the client is created with no arguments, so it reads the key from the environment. The system prompt says British spelling, no medical claims, T B C for missing information. The describe function sends one product and catches rate limits, API errors and connection problems separately. I run it. A sixty word description for the rose hydrating mist appears, and underneath, the usage line, input and output tokens. I note both numbers so I can estimate cost per product.
7:27 Watch me do it, part 2
Next, structured outputs. The schema has sentiment and topic as fixed lists, needs reply as true or false, and a summary, with no extra fields allowed. I pass it in the output configuration and loop over ten anonymised reviews. Each result prints as clean JSON. I paste them next to my own labels in a sheet. Nine agree. One review complains about a courier but was tagged product. So I add a rule to the system prompt, complaints about couriers or delivery times are topic delivery, and run the ten again. Ten out of ten. That little table is now my evaluation set, and I will rerun it whenever I change the prompt or the model.
8:17 Recap and next step
Recap. The API is for software that calls Claude. Master the Messages API, system prompts, model IDs and max tokens. Keep keys in environment variables. Use structured outputs when code needs reliable data. Pull the cost levers, and build an evaluation set before launch. Your next step: run the first call example with your own key on a non sensitive task, log the token usage, then try the review classifier on ten real examples.
When to use the API instead of the app
Use the Claude apps when a person is in the loop for each task. Use the Claude API when software should call Claude: processing every inbound lead, tagging thousands of reviews, powering a support assistant, generating product descriptions from a catalogue feed. You create an account in the Claude Console, generate an API key, and pay per use (billed in tokens, roughly word pieces, with separate prices for input and output that vary by model). Claude is also available through Amazon Bedrock, Google Cloud Vertex AI and Microsoft Foundry if your company already buys cloud services there.
The core concepts
- Messages API: one endpoint (
POST /v1/messages). You send a model name, a list of messages and optional settings; you get content blocks back. - System prompt: sets Claude's role, rules and context for the whole interaction (the API equivalent of Project instructions).
- Model IDs: exact strings such as
claude-sonnet-5,claude-opus-5,claude-haiku-4-5. Check the models overview page for the current list and prices; newer models appear regularly. - max_tokens: the ceiling on output length.
- Thinking and effort: current models support adaptive thinking; an
effortsetting trades depth for speed and cost. - Structured outputs: constrain the response to a JSON schema so your code can rely on it.
- Tool use: you describe functions; Claude asks to call them; your code runs them and returns results.
- Cost levers: prompt caching (reuse a long, stable prefix cheaply), the Message Batches API (asynchronous processing at a discount, currently 50%), and choosing the smallest model that meets your quality bar.
Hands-on: your first call (Python)
pip install anthropic
export ANTHROPIC_API_KEY="sk-ant-..." # from the Claude Console; never commit this
export CLAUDE_MODEL="claude-sonnet-5" # check the models page for current IDsimport os
import anthropic
client = anthropic.Anthropic() # reads ANTHROPIC_API_KEY from the environment
MODEL = os.environ.get("CLAUDE_MODEL", "claude-sonnet-5")
SYSTEM = (
"You write product descriptions for Noor Organics, a UK/UAE skincare brand. "
"British spelling. No medical claims. Never invent ingredients or results; "
"if information is missing, write [TBC]."
)
def describe(product: dict) -> str:
try:
resp = client.messages.create(
model=MODEL,
max_tokens=1000,
system=SYSTEM,
messages=[{
"role": "user",
"content": f"Write a 60-word description for this product:\n{product}",
}],
)
except anthropic.RateLimitError:
raise RuntimeError("Rate limited - retry later or slow down the batch")
except anthropic.APIStatusError as e:
raise RuntimeError(f"Claude API error {e.status_code}: {e.message}")
except anthropic.APIConnectionError:
raise RuntimeError("Network problem reaching the Claude API")
text = "".join(b.text for b in resp.content if b.type == "text")
print("tokens in/out:", resp.usage.input_tokens, resp.usage.output_tokens)
return text
if __name__ == "__main__":
print(describe({"name": "Rose Hydrating Mist", "size": "100ml",
"key_ingredients": ["rose water", "glycerin"]}))The SDK retries some transient errors automatically; the explicit handlers make failures readable. Log usage so you can estimate cost per item.
Structured outputs: JSON your code can trust
import json
schema = {
"type": "object",
"properties": {
"sentiment": {"type": "string", "enum": ["positive", "neutral", "negative"]},
"topic": {"type": "string", "enum": ["delivery", "product", "price", "service", "other"]},
"needs_reply": {"type": "boolean"},
"summary": {"type": "string"},
},
"required": ["sentiment", "topic", "needs_reply", "summary"],
"additionalProperties": False,
}
resp = client.messages.create(
model=MODEL,
max_tokens=500,
messages=[{"role": "user", "content": f"Classify this review:\n{review_text}"}],
output_config={"format": {"type": "json_schema", "schema": schema}},
)
data = json.loads(next(b.text for b in resp.content if b.type == "text"))Quality and safety before launch
- Evaluation set: 30 to 100 real examples with expected outputs; run it on every prompt or model change.
- Guardrails: system-prompt rules, schema constraints, and code checks (for example, block outputs containing banned claims).
- Human hand-off: route low-confidence or sensitive cases (complaints, legal threats, health) to a person.
- Secrets: API keys live in environment variables or a secrets manager, server-side only, never in browser code or repositories. Rotate immediately if exposed.
- Monitoring: log inputs, outputs (minus personal data where possible), latency, cost and error rates. Admins can review usage and costs in the Console.
Worked example: review tagging for a marketplace seller
A Karachi-based electronics seller receives around 2,000 reviews a month across marketplaces. A 70-line script sends each review to a fast model with the schema above, writes results to a Google Sheet, and flags needs_reply items for the support lead. They validated the prompt on 100 hand-labelled reviews first, then moved the monthly back-catalogue to the Batches API to cut cost. Every negative "product" review is still read by a human.
Pitfalls
- Hard-coding keys, or calling the API directly from a public website.
- Skipping the evaluation set and "testing in production".
- Using the most expensive model for simple classification.
- Ignoring
stop_reason(for example, output cut off atmax_tokens).
How to measure success
Accuracy on your evaluation set meets an agreed bar, cost per item is known and acceptable, errors are handled gracefully, and no key has ever been exposed.
Key takeaways
- Use the API when software should call Claude; you pay per token and choose exact model IDs from the current models page.
- Core pieces: Messages API, system prompt, messages, max_tokens, thinking/effort, structured outputs and tool use.
- Keep API keys in environment variables or a secrets manager, server-side only; handle errors and log usage.
- Launch only after an evaluation set, guardrails, human hand-off and monitoring are in place; use caching, batches and right-sized models to control cost.
Check your understanding
Quick questions to lock in the lesson. They don’t count towards your certificate.
Put it into practice
Run the first-call example with your own key on a non-sensitive task and log token usage. Then classify 10 real (anonymised) reviews with the structured-output schema and check each result by hand.
Enrol for free to save your progress
Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.