---
title: "Integrating AI into existing systems | Optimize All Academy"
description: "Where AI meets the rest of your stack An AI feature is only as useful as its connections: the CRM it reads, the help desk it drafts into, the order…"
url: https://optimizeall.com/learn/building-ai-products-and-workflows/integrating-ai-into-existing-systems
updated: 2026-10-05
---

Building AI Products & Workflows · Architecture, models and orchestration · lesson 8 of 18 · 14 min

# Integrating AI into existing systems

## Where AI meets the rest of your stack

An AI feature is only as useful as its connections: the CRM it reads, the help desk it drafts into, the order system it checks, the data warehouse it reports from. Integration is where most of the engineering effort goes, and where reliability problems hide. This lesson covers the integration patterns that work, the failure modes specific to AI calls, and how open standards such as MCP change the picture.

## Integration patterns

| Pattern | How it works | Use when |
|---|---|---|
| Synchronous API call | User action calls your AI service and waits | Interactive features that must answer in seconds |
| Asynchronous job + queue | Event places a job on a queue; a worker processes it and writes back | Anything slow, bulky or retry-prone (summaries, enrichment, document processing) |
| Webhooks in, webhooks out | Source system notifies you of events; you call back when done | SaaS tools (help desks, CRMs, form tools, payment platforms) |
| Scheduled batch | A job runs on a schedule over many records | Reports, nightly enrichment, re-indexing |
| Tools exposed to models | Your systems' capabilities offered as tools, often via MCP servers | Assistants and agents that need to look things up or act |

Most AI features combine several: a webhook triggers a queued job, the worker calls a model that uses tools, and the result is written back through the source system's API.

## Reliability rules for AI integrations

1. **Verify inbound webhooks** (signatures, timestamps) before doing anything, and respond quickly; do the slow work asynchronously.
2. **Idempotency:** the same event may arrive twice. Record processed event IDs and make writes safe to repeat.
3. **Timeouts, retries and back-off:** model APIs return rate-limit and overload errors under load. Retry the retryable ones (rate limits, 5xx, connection errors) with exponential back-off and jitter; do not retry invalid requests.
4. **Dead-letter queues:** after the last retry, park the job with its error for a human, rather than dropping it.
5. **Write back safely:** use the target system's API with least-privilege credentials; prefer creating drafts, notes or suggestions over overwriting records.
6. **Keep AI out of the transaction path** when possible: a failure in the AI step should not block the customer's order or the agent's ticket.
7. **Trace across systems:** carry a correlation ID from the inbound event through the model call to the write-back.

## MCP and the integration layer

MCP changes who builds what. Instead of wiring each system into each AI application, the team that owns a system can publish an MCP server once, and every compatible assistant, IDE or agent platform can use it under the host's consent and permission controls. For internal platforms this often means a small set of **well-designed, least-privilege MCP servers** (CRM, orders, knowledge base) maintained by system owners, plus your own AI services that call them. The same security rules apply: allow-list servers, scope credentials, gate writes and log calls.

## Hands-on: a webhook-to-queue-to-model worker

```python
# app.py  (pip install fastapi uvicorn anthropic)
import hashlib, hmac, json, os, queue, threading, time, random
from fastapi import FastAPI, Request, HTTPException
import anthropic

app = FastAPI()
jobs: "queue.Queue[dict]" = queue.Queue()          # use a real queue (SQS, Pub/Sub, Redis, RabbitMQ) in production
processed: set[str] = set()                         # use a database table in production
SECRET = os.environ["HELPDESK_WEBHOOK_SECRET"].encode()
client = anthropic.Anthropic()
MODEL = os.environ.get("ANTHROPIC_MODEL", "claude-opus-5")

@app.post("/webhooks/helpdesk")
async def helpdesk_webhook(request: Request):
    body = await request.body()
    sig = request.headers.get("X-Signature", "")
    expected = hmac.new(SECRET, body, hashlib.sha256).hexdigest()
    if not hmac.compare_digest(sig, expected):
        raise HTTPException(status_code=401, detail="bad signature")
    event = json.loads(body)
    if event["event_id"] in processed:
        return {"status": "duplicate_ignored"}
    jobs.put(event)                                  # respond fast; work happens in the worker
    return {"status": "queued"}

def draft_with_retries(ticket_text: str, max_attempts: int = 5) -> str:
    for attempt in range(max_attempts):
        try:
            r = client.messages.create(model=MODEL, max_tokens=1200,
                system="Draft a helpful first reply for a human agent to review. Do not promise refunds.",
                messages=[{"role": "user", "content": ticket_text}])
            return "".join(b.text for b in r.content if b.type == "text")
        except (anthropic.RateLimitError, anthropic.InternalServerError, anthropic.APIConnectionError):
            time.sleep(min(60, 2 ** attempt) + random.random())      # exponential back-off with jitter
        except anthropic.BadRequestError:
            raise                                                   # not retryable: fix the request
    raise RuntimeError("model unavailable after retries")

def worker():
    while True:
        event = jobs.get()
        try:
            draft = draft_with_retries(event["ticket"]["text"])
            # write back as an INTERNAL draft via the help desk API (least-privilege token), e.g.:
            # helpdesk.create_draft(ticket_id=event["ticket"]["id"], body=draft, correlation_id=event["event_id"])
            processed.add(event["event_id"])
        except Exception as err:
            print("dead-letter:", event["event_id"], repr(err))       # park for a human; alert the owner

threading.Thread(target=worker, daemon=True).start()
```

Run with `uvicorn app:app`. In production, replace the in-memory queue and set with durable services, add structured logging with the correlation ID, and alert on the dead-letter queue.

## Worked example

A Saudi e-commerce brand connects its help desk to an AI drafting service. Early version: the help desk called the AI synchronously; during a sale-day traffic spike, rate limits caused timeouts and the help desk's webhook retries created duplicate drafts. Fix: verify and queue webhooks, dedupe by event ID, back off on rate limits, write drafts as internal notes, and send failures to a dead-letter queue monitored by the support lead. The next sale day produced no duplicates and no blocked tickets, and the few dead-lettered jobs were handled by hand.

## Video lecture: Integrating AI into existing systems

Lecture coming soon · 15 chapters · about 8 minutes. Read the full transcript below.

1. Integrating AI into your stack
2. Analogy: a new hospital specialist
3. Five patterns
4. Reliability rules 1–4
5. Reliability rules 5–7
6. MCP as an integration layer
7. Simple example: the double webhook
8. Worked example: sale-day spike
9. Business example (illustrative)
10. Hands-on in the lesson
11. Common mistakes
12. How you'll know it's solid
13. Watch me do it: app.py
14. Recap
15. Try this now (20 minutes)

## Lecture transcript

### Integrating AI into your stack

An AI feature that can't read your CRM or write into your help desk is a demo. Integration is where most of the real engineering happens, and where most reliability problems hide: duplicate drafts, timeouts on busy days, jobs that silently vanish. In this lesson you'll learn the integration patterns that work, the reliability rules specific to AI calls, how MCP changes who builds what, and you'll see a webhook-to-queue-to-model worker you can adapt.

### Analogy: a new hospital specialist

Here's an analogy. Integrating AI is like adding a new specialist to a busy hospital. Their skill matters, but so do the handovers: how patients are referred to them, how results get back into the patient record, what happens if they're off sick, and how nobody gets treated twice by mistake. Webhooks, queues, retries and idempotency are the hospital's handover procedures for your AI.

### Five patterns

Five patterns cover most needs. A synchronous API call, where the user waits, for interactive features that must answer in seconds. An asynchronous job on a queue, for anything slow, bulky or retry-prone, like summaries or document processing. Webhooks in and out, for SaaS tools like help desks, CRMs and form builders. Scheduled batches, for reports and nightly enrichment. And tools exposed to models, often through MCP servers, for assistants and agents. Most real features combine several: a webhook triggers a queued job, the worker calls a model with tools, and the result is written back.

### Reliability rules 1–4

Now the reliability rules. Verify inbound webhooks by checking signatures and timestamps, respond quickly, and do slow work asynchronously. Make everything idempotent, because the same event may arrive twice: record processed event IDs. Retry only retryable errors, like rate limits, overloads and connection failures, with exponential back-off and a little randomness. Don't retry invalid requests. After the last retry, park the job in a dead-letter queue for a human, rather than dropping it.

### Reliability rules 5–7

Three more rules. Write back safely, using the target system's API with least-privilege credentials, and prefer drafts, notes or suggestions over overwriting records. Keep AI out of the transaction path, so a failed AI step never blocks a customer's order or an agent's ticket. And trace across systems by carrying one correlation ID from the inbound event, through the model call, to the write-back. When something goes wrong, you'll find it in minutes.

### MCP as an integration layer

MCP changes who builds what. Instead of wiring every system into every AI application, the team that owns a system can publish one MCP server, and every compatible assistant, IDE or agent platform can use it under the host's consent and permission controls. For many organisations that means a small set of well-designed, least-privilege servers for core systems like the CRM, orders and the knowledge base, maintained by their owners. The same security rules still apply: allow-lists, scoped credentials, gated writes and logs.

### Simple example: the double webhook

A simple example of idempotency. Your form tool sends a webhook when someone requests a quote. Occasionally, it sends the same webhook twice. Without protection, your AI drafts two quotes and emails the customer twice. With idempotency, your service records each event ID the first time it's processed. When the duplicate arrives, it sees the ID and does nothing. One line of logic, and a whole class of embarrassing errors disappears.

### Worked example: sale-day spike

Here's a real-world shaped story. A Saudi e-commerce brand connects its help desk to an AI drafting service. At first, the help desk calls the AI synchronously. On a big sale day, rate limits cause timeouts, the help desk retries its webhooks, and agents find duplicate drafts on tickets. The fix: verify and queue webhooks, dedupe by event ID, back off on rate limits, write drafts as internal notes, and send failures to a dead-letter queue watched by the support lead. The next sale day: no duplicates, no blocked tickets, and a handful of parked jobs handled by hand.

### Business example (illustrative)

Illustrative numbers for the sale-day story. On the first sale day, the help desk received roughly ten times normal volume; about one ticket in eight got a duplicate draft and dozens timed out. On the next sale day, with queueing, dedupe and back-off, there were zero duplicates, drafts arrived within a couple of minutes even at peak, and just seven jobs landed in the dead-letter queue, all handled by hand within the hour.

### Hands-on in the lesson

In the hands-on section you'll read a compact FastAPI service. It verifies a help desk webhook with an HMAC signature, ignores duplicate event IDs, queues the job and responds immediately. A worker drafts a reply with the Claude API, retrying rate-limit, server and connection errors with exponential back-off and jitter, never retrying bad requests, and dead-letters anything that still fails. In production, you'd swap the in-memory queue and set for durable services and add structured logs with the correlation ID.

### Common mistakes

Common mistakes. Calling the AI synchronously inside a customer's checkout or ticket submission. Retrying every error, including invalid requests that will never succeed. No dead-letter queue, so failures vanish. Overwriting records directly instead of creating drafts or notes. Using one powerful API key for everything. And no correlation ID, so tracing a problem across three systems takes a day.

### How you'll know it's solid

How will you know your integration is solid? Duplicates never produce duplicate actions. Rate-limit spikes cause delays, not failures. Every failed job ends up somewhere a person can see it. Core business flows keep working when the AI provider is down. And you can follow one customer's request from the inbound event to the written-back result in a single search, using its correlation ID.

### Watch me do it: app.py

Watch me do it. I open app dot py. First, the webhook route reads the raw body and computes an HMAC with our shared secret. If it doesn't match the signature header, it returns four-oh-one. Next, it parses the event and checks the event ID against the processed set; a duplicate returns duplicate ignored. Otherwise it puts the event on the queue and returns queued immediately. Then the worker. It takes a job and calls draft with retries. That function catches rate-limit, server and connection errors and sleeps with exponential back-off plus jitter, but re-raises bad request errors, because retrying won't fix them. On success, the worker would write an internal draft back to the help desk with the event ID as correlation ID, and marks the event processed. On failure, it prints a dead-letter line. I send the same test webhook twice. The first returns queued, the second duplicate ignored.

### Recap

To recap: pick integration patterns deliberately and combine them. Verify, dedupe, retry sensibly, dead-letter, write back with least privilege, keep AI off the transaction path and trace end to end. Use MCP to publish system integrations once, securely. Your next step is to draw the integration for one AI feature, marking where idempotency, retries and dead-lettering happen. Next module: data, privacy and security architecture.

### Try this now (20 minutes)

Try this now. Pick one AI feature and draw its integration: the trigger, whether it's synchronous or queued, the model call, any tools, the write-back and the failure path. Mark where idempotency, retries and dead-lettering happen, and which credentials each step uses. If any of those marks are missing, that's your first fix, before you add any new AI capability.

## Key takeaways

- Choose integration patterns deliberately: synchronous calls, queued jobs, webhooks, scheduled batches and tools/MCP.
- Verify webhooks, dedupe by event ID, retry only retryable errors with back-off and jitter, and dead-letter failures.
- Write back with least privilege, prefer drafts and notes, and keep AI out of the transaction path.
- MCP lets system owners publish integrations once for many AI apps, under the same security controls.

## Try it

Draw the integration for one AI feature: trigger, queue or sync call, model call, tools, write-back and failure path. Mark where idempotency, retries and dead-lettering happen.

- [Previous: Workflow orchestration patterns for AI features](https://optimizeall.com/learn/building-ai-products-and-workflows/workflow-orchestration-patterns)
- [Next: Data and privacy architecture for AI features](https://optimizeall.com/learn/building-ai-products-and-workflows/data-and-privacy-architecture)
- [All lessons of Building AI Products & Workflows](https://optimizeall.com/learn/building-ai-products-and-workflows)
