Skip to content

AI Automation with n8n, Make and Zapier · Reliability, safety and cost · lesson 11 of 17 · 17 min

Error handling, retries and idempotency

Everything fails eventually

APIs time out, tokens expire, rate limits hit, apps change fields, and LLMs return malformed output. A workflow without error handling is not finished; it is waiting to fail silently. Reliability engineering for automations rests on four ideas: classify errors, retry the right ones, make actions idempotent, and alert humans with context.

Classify errors

| Class | Examples | Response | |---|---|---| | Transient | Timeouts, 429 rate limits, 5xx server errors, network blips | Retry with exponential backoff and jitter | | Data/validation | 400/422, missing required field, invalid email, malformed AI JSON | Do not blind-retry; route to a fix queue or human; alert | | Auth/config | 401/403, expired OAuth token, revoked permission | Stop, alert the owner immediately; reconnect credential | | Logic/bugs | Unexpected nulls, wrong mapping after an app change | Alert, fix the workflow, replay affected runs |

Retries done right

  • Exponential backoff with jitter: wait 2s, 4s, 8s... plus randomness, up to a cap.
  • Limit attempts (for example 3 to 5), then escalate.
  • Retry only idempotent or idempotency-keyed operations; otherwise you risk duplicates.

Platform features:

  • n8n: node settings "Retry On Fail" (max tries, wait between tries); "On Error" behavior (stop, continue, or continue using an error output branch); a separate Error Workflow (using the Error Trigger node) that runs when a workflow fails, ideal for alerts.
  • Make: error handler routes on modules: Break (store as incomplete execution, with optional automatic retries), Resume (provide fallback output), Ignore, Commit, Rollback; plus incomplete executions to fix and re-run.
  • Zapier: Autoreplay (plan-dependent) retries failed steps; error notifications; Paths and Filters for fallbacks; custom error handling where available.

Idempotency: safe to run twice

An operation is idempotent if running it twice has the same effect as once. Techniques:

  1. Upserts (create or update by a unique key such as email or order ID) instead of blind creates.
  2. Idempotency keys: many APIs (for example payment APIs) accept an Idempotency-Key header; the server ignores repeats with the same key.
  3. Processed-ID stores: record event or order IDs in a data store, table or database; skip if already processed.
  4. State checks: "only send reminder if reminded_at is empty", then set it.

Dead-letter and replay

Failed items should land somewhere visible (a "dead-letter" table, Slack channel or helpdesk queue) with the payload, error and a link to the execution. After fixing, replay: re-run from the failed step (n8n executions, Make incomplete executions, Zapier replay).

Monitoring

  • Alert on failure, and on silence: a daily workflow that did not run is also an incident (heartbeat checks).
  • Track run counts, error rates, and duration.
  • Review weekly: top errors, root causes, fixes.

Worked example: a UAE clinic's booking sync

A Make scenario syncs new website bookings to the clinic's practice-management API and sends confirmations.

  • Practice API sometimes returns 503 at peak hours: Break handler with automatic retries (3 attempts, increasing intervals).
  • Invalid phone formats return 422: Resume is wrong here; instead route to a Google Sheet "fix queue" plus a Slack alert to reception.
  • Confirmation messages use a data store of booking IDs to prevent duplicate WhatsApp messages on retries.
  • A daily heartbeat scenario counts yesterday's bookings; if zero on a weekday, it alerts the operations manager.

Hands-on: an n8n error workflow

  1. Create a workflow "Global error handler" starting with Error Trigger.
  2. Add a Code node to format a message:
const e = $json;
return [{ json: {
  text: `:rotating_light: ${e.workflow.name} failed at "${e.execution.lastNodeExecuted}"\n` +
        `Error: ${e.execution.error.message}\nExecution: ${e.execution.url}`
}}];
  1. Add a Slack (or email) node posting to #automation-alerts.
  2. In each production workflow's settings, set Error Workflow to "Global error handler".
  3. On HTTP nodes, enable Retry On Fail (for example 3 tries, 5,000 ms wait) for transient calls only, and use an error output branch for 4xx data errors.

(Check the Error Trigger's output fields in your n8n version; field names can differ.)

Pitfalls

  • Retrying 4xx validation errors forever.
  • Retrying non-idempotent creates, producing duplicates.
  • Alerts with no context ("Workflow failed") that nobody can act on.

Measuring success

  • Mean time to detect a failure (alerts should arrive within minutes).
  • Mean time to recover (fix plus replay).
  • Duplicate side effects per month (target: zero).
  • Share of failures that were transient and self-healed by retries versus needing a human.

Video lecture: Error handling, retries and idempotency

Lecture coming soon · 13 chapters · about 8 minutes. Read the full transcript below.

  1. Error handling
  2. Why it matters
  3. Analogy: an airline
  4. Classify errors
  5. Platform tools
  6. Idempotency
  7. Example 1: n8n global error handler
  8. Example 2: UAE clinic booking sync
  9. Dead letters + monitoring
  10. Common mistakes
  11. Watch me do it: make it robust
  12. Recap + try this now
  13. Try this now

Lecture transcript

Error handling

Here's a hard truth about automation. Every workflow you build will fail at some point. APIs time out, tokens expire, rate limits hit, apps rename fields, and language models occasionally return broken JSON. The question isn't whether it fails. It's whether it fails loudly, safely and recoverably, or silently, creating duplicates and angry customers. In this lesson you'll learn to classify errors, retry correctly, make actions idempotent, catch failures in dead-letter queues, and monitor your automations.

Why it matters

Why does this matter? Because silent failures are expensive. A lead-routing workflow that quietly stops on a Friday means no follow-ups all weekend. A booking sync that retries a create call can send three confirmation messages to one patient. And an expired token can stop everything until someone notices. Robust automations are what separate a hobby project from something a business can depend on.

Analogy: an airline

An analogy. Think of an airline. Delays happen, so they have rules. Short weather delays? Wait and try again. That's a retry. A passenger with the wrong passport? No amount of waiting fixes it. That goes to a desk for a human. That's a data error. A pilot without a license? Ground the plane and call the manager. That's an auth or config error. And every booking has a reference number, so rebooking twice doesn't give you two seats. That's idempotency.

Classify errors

So, classify errors. Transient errors, like timeouts, four twenty-nine rate limits and five hundreds, get retried with exponential backoff and jitter: wait two seconds, then four, then eight, plus a little randomness, up to a limit. Data errors, like four hundreds, four twenty-twos, missing fields or malformed AI output, should not be blindly retried. Send them to a fix queue and alert someone. Auth and config errors, like four oh ones and expired tokens, should stop the workflow and alert the owner immediately. And logic bugs get fixed, then the affected runs replayed.

Platform tools

Each platform has tools for this. In n8n, node settings offer retry on fail with a maximum number of tries and a wait, plus an on error setting that can continue through an error output branch, and an error workflow, started by the error trigger, that runs whenever a workflow fails. In Make, you attach error handler routes: break stores the run as incomplete with optional automatic retries, resume substitutes fallback output, ignore skips, and commit and rollback handle transactions. In Zapier, autoreplay retries failed steps on eligible plans, and error notifications tell you what broke.

Idempotency

Now idempotency: making it safe to run twice. Use upserts, create or update by a unique key, instead of blind creates. Use idempotency keys, a header many APIs accept, so repeats with the same key are ignored. Keep a store of processed IDs, and skip anything already handled. And check state before acting: only send a reminder if reminded at is empty, then set it. With idempotency in place, retries become safe instead of dangerous.

Example 1: n8n global error handler

Example one, simple. An n8n global error handler. One workflow starts with the error trigger, formats a message with the workflow name, the failing node, the error message and a link to the execution, and posts it to an automation alerts Slack channel. Every production workflow points to it in settings. HTTP nodes get retry on fail for transient calls, and an error output branch for data errors. The code is in the lesson text. Ten minutes of setup, and you'll never have a silent failure again.

Example 2: UAE clinic booking sync

Example two, realistic. A clinic in the UAE uses Make to sync website bookings to its practice management system. At peak hours the API sometimes returns five oh three, so a break handler retries three times with growing intervals. Invalid phone numbers return four twenty-two, and those go to a fix queue in Google Sheets with a Slack alert to reception, not retries. Confirmations check a data store of booking IDs so retries never send duplicate WhatsApp messages. And a daily heartbeat scenario alerts the operations manager if a weekday shows zero bookings, because silence can be a failure too.

Dead letters + monitoring

Dead letters and replay tie it together. Failed items should land somewhere visible, like a table, Slack channel or helpdesk queue, with the payload, the error and a link to the execution. After you fix the cause, replay from the failed step. And monitor: alert on failures and on silence, track run counts, error rates and durations, and review the top errors weekly.

Common mistakes

Three mistakes cause most pain. Retrying four hundred validation errors forever, which never succeeds and floods logs. Retrying non-idempotent creates, which produces duplicates. And alerts with no context, just workflow failed, which nobody can act on. Every alert should say what failed, where, why, and link to the run.

Watch me do it: make it robust

Watch me do it. I make our n8n lead workflow robust. First, the global error handler: a new workflow with an Error Trigger, a Code node that formats the workflow name, the failing node, the error message and the execution link, and a Slack node posting to automation alerts. In the settings of both lead workflows, I select it as the error workflow. Second, retries: on the CRM and courier HTTP nodes I enable retry on fail for transient errors, three tries with a five-second wait. For data errors I use error outputs instead, routed to a Google Sheet called fix queue plus a Slack message. Third, idempotency: at the start of process lead I add a lookup in a data table of processed lead IDs, built from the form submission ID. If the ID exists, the workflow stops quietly and logs a duplicate. At the end, it writes the ID. The CRM step uses create or update by email rather than create. Fourth, a heartbeat: a scheduled workflow every weekday at ten in the morning, Karachi time, counts yesterday's leads and alerts if the number is zero. Now the test: I copy the workflow, break the CRM credential in the copy, and run it. Within seconds an alert arrives in Slack with the workflow name, the failing node and a link to the execution.

Recap + try this now

Recap. Classify errors: retry transient ones with backoff, route data errors to humans, stop on auth problems. Make actions idempotent with upserts, keys, processed-ID stores and state checks. Send failures to a visible dead-letter queue and replay after fixing. Alert on failure and silence. Try this now: set up a global error handler on your platform, point every production workflow at it, then deliberately break one API credential in a test copy and confirm the alert arrives with enough context to act.

Try this now

Try this now. On your platform, create a global error handler: an n8n error workflow, a Make error route pattern, or Zapier error notifications to a shared channel. Point every production workflow at it. On one HTTP step, enable retries for transient errors only. Add a processed-ID check to one workflow that creates records. Then make a test copy of a workflow, break its credential on purpose, run it, and confirm the alert arrives with the workflow name, failing step, error and a link.

Key takeaways

  • Classify errors: retry transient ones (timeouts, 429, 5xx) with backoff and jitter; route data errors to humans; stop and alert on auth/config errors.
  • Use platform features: n8n Retry On Fail, error outputs and Error Workflows; Make Break/Resume/Ignore/Commit/Rollback; Zapier Autoreplay and notifications.
  • Make actions idempotent with upserts, idempotency keys, processed-ID stores and state checks so retries are safe.
  • Send failures to a visible dead-letter queue with context, replay after fixes, and alert on both failures and silence.

Try it

Set up a global error handler (n8n Error Workflow, Make error routes or Zapier notifications). Link your production workflows, then break a credential in a test copy and confirm the alert has enough context to act.