Mastering Claude (Anthropic)Building with the Claude API · Lesson 16 of 20

The Claude API: core concepts and your first working integration

Article · 18 min · 9 min lecture

Video lecture

The Claude API: core concepts and your first working integration

14 chapters · about 9 min · full transcript

Coming soon

Chapter 1 of 14

The Claude API

  • When software calls Claude
  • Your first working integration

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

When to use the API instead of the app

Use the Claude apps when a person is in the loop for each task. Use the Claude API when software should call Claude: processing every inbound lead, tagging thousands of reviews, powering a support assistant, generating product descriptions from a catalogue feed. You create an account in the Claude Console, generate an API key, and pay per use (billed in tokens, roughly word pieces, with separate prices for input and output that vary by model). Claude is also available through Amazon Bedrock, Google Cloud Vertex AI and Microsoft Foundry if your company already buys cloud services there.

The core concepts

  • Messages API: one endpoint (POST /v1/messages). You send a model name, a list of messages and optional settings; you get content blocks back.
  • System prompt: sets Claude's role, rules and context for the whole interaction (the API equivalent of Project instructions).
  • Model IDs: exact strings such as claude-sonnet-5, claude-opus-5, claude-haiku-4-5. Check the models overview page for the current list and prices; newer models appear regularly.
  • max_tokens: the ceiling on output length.
  • Thinking and effort: current models support adaptive thinking; an effort setting trades depth for speed and cost.
  • Structured outputs: constrain the response to a JSON schema so your code can rely on it.
  • Tool use: you describe functions; Claude asks to call them; your code runs them and returns results.
  • Cost levers: prompt caching (reuse a long, stable prefix cheaply), the Message Batches API (asynchronous processing at a discount, currently 50%), and choosing the smallest model that meets your quality bar.

Hands-on: your first call (Python)

pip install anthropic
export ANTHROPIC_API_KEY="sk-ant-..."   # from the Claude Console; never commit this
export CLAUDE_MODEL="claude-sonnet-5"    # check the models page for current IDs
import os
import anthropic

client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY from the environment
MODEL = os.environ.get("CLAUDE_MODEL", "claude-sonnet-5")

SYSTEM = (
    "You write product descriptions for Noor Organics, a UK/UAE skincare brand. "
    "British spelling. No medical claims. Never invent ingredients or results; "
    "if information is missing, write [TBC]."
)

def describe(product: dict) -> str:
    try:
        resp = client.messages.create(
            model=MODEL,
            max_tokens=1000,
            system=SYSTEM,
            messages=[{
                "role": "user",
                "content": f"Write a 60-word description for this product:\n{product}",
            }],
        )
    except anthropic.RateLimitError:
        raise RuntimeError("Rate limited - retry later or slow down the batch")
    except anthropic.APIStatusError as e:
        raise RuntimeError(f"Claude API error {e.status_code}: {e.message}")
    except anthropic.APIConnectionError:
        raise RuntimeError("Network problem reaching the Claude API")

    text = "".join(b.text for b in resp.content if b.type == "text")
    print("tokens in/out:", resp.usage.input_tokens, resp.usage.output_tokens)
    return text

if __name__ == "__main__":
    print(describe({"name": "Rose Hydrating Mist", "size": "100ml",
                    "key_ingredients": ["rose water", "glycerin"]}))

The SDK retries some transient errors automatically; the explicit handlers make failures readable. Log usage so you can estimate cost per item.

Structured outputs: JSON your code can trust

import json

schema = {
    "type": "object",
    "properties": {
        "sentiment": {"type": "string", "enum": ["positive", "neutral", "negative"]},
        "topic": {"type": "string", "enum": ["delivery", "product", "price", "service", "other"]},
        "needs_reply": {"type": "boolean"},
        "summary": {"type": "string"},
    },
    "required": ["sentiment", "topic", "needs_reply", "summary"],
    "additionalProperties": False,
}

resp = client.messages.create(
    model=MODEL,
    max_tokens=500,
    messages=[{"role": "user", "content": f"Classify this review:\n{review_text}"}],
    output_config={"format": {"type": "json_schema", "schema": schema}},
)
data = json.loads(next(b.text for b in resp.content if b.type == "text"))

Quality and safety before launch

  1. Evaluation set: 30 to 100 real examples with expected outputs; run it on every prompt or model change.
  2. Guardrails: system-prompt rules, schema constraints, and code checks (for example, block outputs containing banned claims).
  3. Human hand-off: route low-confidence or sensitive cases (complaints, legal threats, health) to a person.
  4. Secrets: API keys live in environment variables or a secrets manager, server-side only, never in browser code or repositories. Rotate immediately if exposed.
  5. Monitoring: log inputs, outputs (minus personal data where possible), latency, cost and error rates. Admins can review usage and costs in the Console.

Worked example: review tagging for a marketplace seller

A Karachi-based electronics seller receives around 2,000 reviews a month across marketplaces. A 70-line script sends each review to a fast model with the schema above, writes results to a Google Sheet, and flags needs_reply items for the support lead. They validated the prompt on 100 hand-labelled reviews first, then moved the monthly back-catalogue to the Batches API to cut cost. Every negative "product" review is still read by a human.

Pitfalls

  • Hard-coding keys, or calling the API directly from a public website.
  • Skipping the evaluation set and "testing in production".
  • Using the most expensive model for simple classification.
  • Ignoring stop_reason (for example, output cut off at max_tokens).

How to measure success

Accuracy on your evaluation set meets an agreed bar, cost per item is known and acceptable, errors are handled gracefully, and no key has ever been exposed.

Key takeaways

  • Use the API when software should call Claude; you pay per token and choose exact model IDs from the current models page.
  • Core pieces: Messages API, system prompt, messages, max_tokens, thinking/effort, structured outputs and tool use.
  • Keep API keys in environment variables or a secrets manager, server-side only; handle errors and log usage.
  • Launch only after an evaluation set, guardrails, human hand-off and monitoring are in place; use caching, batches and right-sized models to control cost.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. In the Claude API, what does the system prompt do?
  2. Your code must write Claude's review classification straight into a spreadsheet. What is the most reliable approach?
  3. A colleague pastes an API key into a shared chat. What should happen?

Put it into practice

Run the first-call example with your own key on a non-sensitive task and log token usage. Then classify 10 real (anonymised) reviews with the structured-output schema and check each result by hand.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.