Building AI Products & WorkflowsArchitecture, models and orchestration · Lesson 8 of 18

Integrating AI into existing systems

Article · 14 min · 8 min lecture

Video lecture

Integrating AI into existing systems

15 chapters · about 8 min · full transcript

Coming soon

Chapter 1 of 15

Integrating AI into your stack

  • Patterns
  • Reliability rules
  • MCP's role
  • A worked worker

The narrated lecture is in production

Every chapter is scripted and ready. Browse the chapters and read the full transcript now — the video will appear here when it’s published.

Chapters

Where AI meets the rest of your stack

An AI feature is only as useful as its connections: the CRM it reads, the help desk it drafts into, the order system it checks, the data warehouse it reports from. Integration is where most of the engineering effort goes, and where reliability problems hide. This lesson covers the integration patterns that work, the failure modes specific to AI calls, and how open standards such as MCP change the picture.

Integration patterns

PatternHow it worksUse when
Synchronous API callUser action calls your AI service and waitsInteractive features that must answer in seconds
Asynchronous job + queueEvent places a job on a queue; a worker processes it and writes backAnything slow, bulky or retry-prone (summaries, enrichment, document processing)
Webhooks in, webhooks outSource system notifies you of events; you call back when doneSaaS tools (help desks, CRMs, form tools, payment platforms)
Scheduled batchA job runs on a schedule over many recordsReports, nightly enrichment, re-indexing
Tools exposed to modelsYour systems' capabilities offered as tools, often via MCP serversAssistants and agents that need to look things up or act

Most AI features combine several: a webhook triggers a queued job, the worker calls a model that uses tools, and the result is written back through the source system's API.

Reliability rules for AI integrations

  1. Verify inbound webhooks (signatures, timestamps) before doing anything, and respond quickly; do the slow work asynchronously.
  2. Idempotency: the same event may arrive twice. Record processed event IDs and make writes safe to repeat.
  3. Timeouts, retries and back-off: model APIs return rate-limit and overload errors under load. Retry the retryable ones (rate limits, 5xx, connection errors) with exponential back-off and jitter; do not retry invalid requests.
  4. Dead-letter queues: after the last retry, park the job with its error for a human, rather than dropping it.
  5. Write back safely: use the target system's API with least-privilege credentials; prefer creating drafts, notes or suggestions over overwriting records.
  6. Keep AI out of the transaction path when possible: a failure in the AI step should not block the customer's order or the agent's ticket.
  7. Trace across systems: carry a correlation ID from the inbound event through the model call to the write-back.

MCP and the integration layer

MCP changes who builds what. Instead of wiring each system into each AI application, the team that owns a system can publish an MCP server once, and every compatible assistant, IDE or agent platform can use it under the host's consent and permission controls. For internal platforms this often means a small set of well-designed, least-privilege MCP servers (CRM, orders, knowledge base) maintained by system owners, plus your own AI services that call them. The same security rules apply: allow-list servers, scope credentials, gate writes and log calls.

Hands-on: a webhook-to-queue-to-model worker

# app.py  (pip install fastapi uvicorn anthropic)
import hashlib, hmac, json, os, queue, threading, time, random
from fastapi import FastAPI, Request, HTTPException
import anthropic

app = FastAPI()
jobs: "queue.Queue[dict]" = queue.Queue()          # use a real queue (SQS, Pub/Sub, Redis, RabbitMQ) in production
processed: set[str] = set()                         # use a database table in production
SECRET = os.environ["HELPDESK_WEBHOOK_SECRET"].encode()
client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")

@app.post("/webhooks/helpdesk")
async def helpdesk_webhook(request: Request):
    body = await request.body()
    sig = request.headers.get("X-Signature", "")
    expected = hmac.new(SECRET, body, hashlib.sha256).hexdigest()
    if not hmac.compare_digest(sig, expected):
        raise HTTPException(status_code=401, detail="bad signature")
    event = json.loads(body)
    if event["event_id"] in processed:
        return {"status": "duplicate_ignored"}
    jobs.put(event)                                  # respond fast; work happens in the worker
    return {"status": "queued"}

def draft_with_retries(ticket_text: str, max_attempts: int = 5) -> str:
    for attempt in range(max_attempts):
        try:
            r = client.messages.create(model=MODEL, max_tokens=1200,
                system="Draft a helpful first reply for a human agent to review. Do not promise refunds.",
                messages=[{"role": "user", "content": ticket_text}])
            return "".join(b.text for b in r.content if b.type == "text")
        except (anthropic.RateLimitError, anthropic.InternalServerError, anthropic.APIConnectionError):
            time.sleep(min(60, 2 ** attempt) + random.random())      # exponential back-off with jitter
        except anthropic.BadRequestError:
            raise                                                   # not retryable: fix the request
    raise RuntimeError("model unavailable after retries")

def worker():
    while True:
        event = jobs.get()
        try:
            draft = draft_with_retries(event["ticket"]["text"])
            # write back as an INTERNAL draft via the help desk API (least-privilege token), e.g.:
            # helpdesk.create_draft(ticket_id=event["ticket"]["id"], body=draft, correlation_id=event["event_id"])
            processed.add(event["event_id"])
        except Exception as err:
            print("dead-letter:", event["event_id"], repr(err))       # park for a human; alert the owner

threading.Thread(target=worker, daemon=True).start()

Run with uvicorn app:app. In production, replace the in-memory queue and set with durable services, add structured logging with the correlation ID, and alert on the dead-letter queue.

Worked example

A Saudi e-commerce brand connects its help desk to an AI drafting service. Early version: the help desk called the AI synchronously; during a sale-day traffic spike, rate limits caused timeouts and the help desk's webhook retries created duplicate drafts. Fix: verify and queue webhooks, dedupe by event ID, back off on rate limits, write drafts as internal notes, and send failures to a dead-letter queue monitored by the support lead. The next sale day produced no duplicates and no blocked tickets, and the few dead-lettered jobs were handled by hand.

Key takeaways

  • Choose integration patterns deliberately: synchronous calls, queued jobs, webhooks, scheduled batches and tools/MCP.
  • Verify webhooks, dedupe by event ID, retry only retryable errors with back-off and jitter, and dead-letter failures.
  • Write back with least privilege, prefer drafts and notes, and keep AI out of the transaction path.
  • MCP lets system owners publish integrations once for many AI apps, under the same security controls.

Check your understanding

Quick questions to lock in the lesson. They don’t count towards your certificate.

  1. A help desk retries a webhook and your service creates two AI drafts for one ticket. What prevents this?
  2. Which model API errors should generally be retried with back-off?
  3. Why keep the AI step out of the transaction path where possible?

Put it into practice

Draw the integration for one AI feature: trigger, queue or sync call, model call, tools, write-back and failure path. Mark where idempotency, retries and dead-lettering happen.

Enrol for free to save your progress

Reading is always free. Enrol to keep your place, take the final assessment and earn a verifiable certificate.