A workflow that fails at three in the morning, and tells you so
Everything you have built so far runs when you press the button and you watch it work. The moment a workflow goes Active on a schedule, that stops being true — it runs while you are asleep, and when it dies it dies quietly. This part strings a net under it: an error workflow that messages you when a run gives up, a node-level retry that rescues the failures that were only ever momentary, and the Wait node, which stops you causing the failure in the first place.
Three ways a clever workflow lets you down
By now you can switch a workflow from Manual to Schedule so it runs by itself at 8:00 every morning, and you can split, merge and batch the data going through it. The workflows look clever. Here is what clever looks like on a bad week.
n8n has a ready answer for each one. An error workflow covers the first, retry covers the third, and the Wait node covers the second. They are not alternatives you pick between. They are three layers of insurance that a production workflow should carry at the same time: tell you it broke, let a single point rescue itself, and avoid hammering the other side to begin with.
Three names that look alike and work at completely different levels
Lay the three cards on the table before you touch anything, because beginners mix them up constantly. Retry means this one node tries again by itself. An error workflow means the whole workflow has already given up. Five failed retries → the workflow dies → only then is the error workflow called.
| Mechanism | Which level of failure it catches | What it does | When to use it |
|---|---|---|---|
| Error workflow | The whole workflow dies — any node fails for good | Triggers another workflow built to handle errors, which messages you on Slack, for example | Every production workflow should be bound to one |
| Retry on fail, at node level | One node fails — an HTTP Request dies, say |
That node retries itself a set number of times, with a gap between tries | An API that acts up now and then, a network blip, mild rate limiting |
The Wait node |
It does not catch failures. It prevents them | The workflow stops here for a number of seconds, until a given time, or until an external callback | Avoiding a rate limit from calling too fast, spacing out notifications, waiting for another system to finish |
Think of a kitchen. Not overloading one socket is the Wait node — you never trip the fuse in the first place. The fuse that resets itself after a moment is retry. The smoke alarm that wakes the whole house is the error workflow. Take any one of the three away and the other two still work; you have just lost a line of defense.
The order they fire in matters, and it is worth having the picture in your head before you build anything.
flowchart TD
A["A node fails mid-run"] --> B{"Is Retry On Fail turned on
for this node?"}
B -->|"no"| D
B -->|"yes"| C{"Are there tries left?
Max Tries counts the first run"}
C -->|"yes"| C1["Pause for Wait Between Tries,
then call the node again"]
C1 --> C
C -->|"tries used up"| D["The workflow dies"]
D --> E{"Does this workflow have an
Error Workflow set, and is
that handler Active?"}
E -->|"no"| F["Nobody is told.
You find out on Monday"]
E -->|"yes"| G["The Error Trigger fires
inside Error Handler"]
G --> H["A message names the workflow,
the node and the error"]
One error workflow for the whole company
An error workflow is a separate, standalone workflow. Not a setting — a whole workflow, sitting in your Workflows list like any other. The one thing that makes it special is that its first node has to be an Error Trigger.
It works in three moves. Build the error workflow first, with an Error Trigger at the front and Slack, email, Discord or a database behind it — that answers “who gets told when something breaks, and where does it get logged?” Point every other workflow at it in their own Settings. Then, when one of them fails, n8n triggers the error workflow automatically and hands it the failure data.
sequenceDiagram participant S as Schedule Trigger participant W as Daily report workflow participant N as n8n participant E as Error Handler participant K as Slack S->>W: the 8 am run starts on its own W->>W: the HTTP Request node fails W->>W: retries, and fails again W-->>N: no tries left, the workflow dies N->>E: call the workflow named in Error Workflow E->>E: Error Trigger receives the failure data E->>K: post to the alerts channel K-->>N: you know within seconds, not on Monday
The Error Trigger node puts the error information into $json. These are the five fields you will use most:
| Field | What it means | Expression |
|---|---|---|
| The failed workflow's name | Which workflow died | {{ $json.workflow.name }} |
| The failed workflow's ID | Handy for linking straight into it | {{ $json.workflow.id }} |
| The error message | n8n's own description of the error | {{ $json.execution.error.message }} |
| Which node it died at | The node that failed | {{ $json.execution.lastNodeExecuted }} |
| Execution URL | Click straight through to that execution record | {{ $json.execution.url }} |
Error Trigger node fires only because another automatically triggered workflow died — Schedule, Webhook or an Event Trigger. Running the error workflow yourself with Execute Workflow, or failing a parent workflow you started by hand, does not start it. The docs say it outright: “You can't test error workflows when running workflows manually. The Error Trigger only runs when an automatic workflow errors.” To test it, take a workflow that is certain to fail, set it Active with a Schedule, and let it run and die on its own.It is the shop alarm. You can stand inside during opening hours pushing the back door as hard as you like and nothing rings, because the system is in service mode with you in the building. It only sounds when the shop is locked up and running without you. Annoying when you want to check it works; exactly right the rest of the time.
Error Trigger, if it is Active and gets triggered automatically itself, calls itself when it fails. Do not put a Schedule or Webhook Trigger into your Error Handler — that creates a confusing self-reference. Leave the Error Trigger as its only trigger.Building the Error Handler
The goal: a workflow that posts a message to the Slack #alerts channel whenever something fails. Set it up once, and every workflow you write afterwards uses it.
-
Step 1
Create a new workflow and name it Error Handler
Go to Workflows in the left sidebar → + Add workflow. With the canvas still blank, press
Ctrl+Sto save and type the nameError Handler. Pick something recognizable — you will be looking for it in a dropdown later, and a workflow called “My workflow 7” is no help there. -
Step 2
Add the Error Trigger node
Press
Tabon an empty part of the canvas, or use the + in the top left, to open the nodes panel → search forError Trigger→ pick it. It lands on the canvas as this workflow's starting point. TheError Triggerhas no parameters to set. Dropping it in is all you do. -
Step 3
Add a Slack node behind it
Drag a connection out of the
Error Trigger's right side and let go to open the nodes panel → search forSlack→ pick the Send a message action. If you have not set up a Slack credential yet, create one now — an OAuth authorization for your Slack workspace. -
Step 4
Set the Channel
The Channel field lets you pick an existing channel By Name, or you can type
#alertsstraight in. Create that channel in Slack first and invite the n8n bot into it, from the Slack channel → Integrations → Add App menu. A message sent to a channel the bot is not in fails, and that failure has nothing above it to catch it. -
Step 5
Write the message with expressions
Switch the Text field to Expression mode — the
fxbutton at the top right of the field — and paste in a body built from the fields above. The source guide's version opens with a warning-sign line, then these five, one per line:Workflow: {{ $json.workflow.name }}Failed at node: {{ $json.execution.lastNodeExecuted }}Error message: {{ $json.execution.error.message }}Time: {{ $json.execution.startedAt }}Open it here: {{ $json.execution.url }}The preview underneath the field shows a sample filled with dummy data. If that reads like a message you would want at three in the morning, you are done with this step.
-
Step 6
Save, then flip Active
Press
Ctrl+S, then turn the Active toggle in the top right green. An error workflow that is not Active never fires. This is the number-one gotcha in the whole chapter. -
Step 7
Leave it alone
Once the Error Handler is set up there is nothing to tend day to day. It is called and runs once, only when another workflow fails, and takes no resources at all the rest of the time.
IF node to judge severity, so that billing workflows also place a phone call when they fail while everything else only posts to Slack. This is your central error desk; add whatever you like to it once the simple version is proven.Pointing every workflow at it
The Error Handler existing does nothing on its own. You still have to go into each production workflow and point it at the handler, one at a time. It is not applied automatically.
-
Step 1
Open the workflow you want watched
Click into it from the Workflows list and the canvas opens.
-
Step 2
Three-dot menu, top right → Settings
Or press
Ctrl+,. The Workflow Settings panel opens, with options such as Timezone, Save Failed Executions and Error Workflow. Your build may lay this panel out differently — go by the field names rather than their position. -
Step 3
Pick Error Handler in the Error Workflow dropdown
The dropdown lists every workflow on your instance that contains an
Error Triggernode — and only those. FindError Handlerand click it. -
Step 4
Save Settings
Save sits at the bottom right of the panel. From now on this workflow triggers Error Handler automatically when it fails.
-
Step 5
Repeat for every production workflow
Make a list and work through it in one sitting. Then set it on new workflows as you create them, while you are already in the Settings panel choosing a timezone.
Letting one node rescue itself
Retry on fail is an automatic retry set on a single node. When that node fails, n8n waits a moment and calls it again, waits and calls again, and only declares failure once the tries run out. It suits transient errors: an API returning 503 for a while, a network wobble, the other system restarting.
To set it, open the panel of the node you want to retry — HTTP Request, Slack, Google Sheets, any node can have it. The top of the node panel has two tabs, Parameters and Settings. Click Settings, then turn Retry On Fail on. Two fields appear.
3 means at most three attempts: the first fails, wait, the second fails, wait, and only the third failure is fatal.1000 (1 second) to 5000 (5 seconds). Too short and the other system gets no breathing room; too long and the workflow drags. For rate-limit errors, use 5000 or more.| Scenario | Good candidate for retry | Why |
|---|---|---|
| 503 Service Unavailable | Ideal | Their service is down for a moment and usually comes back within seconds |
| 429 Too Many Requests | Yes, with a longer Wait Between Tries | The block lifts after a wait |
| Network timeout or connection failure | Yes | A network blip is usually over in an instant |
| 401 Unauthorized, the credential expired | No | It is just as expired on the hundredth call. Go and re-authorize the credential |
| 404 Not Found, the data really is not there | No | The data does not exist, and retrying will not conjure it up |
You wrote the expression wrong, and get undefined | No | A logic error does not turn correct because you repeat it |
It is the card machine that says the payment failed. You tap the card again, and the customer gets charged twice, because the first one actually went through. “Try it again” is only safe advice when trying it again is free.
flowchart TD A["A node keeps failing.
Should it retry?"] --> B{"Would running this node twice
send, create or charge twice?"} B -->|"yes"| B1["Leave Retry On Fail off.
Put a check-first node in front,
or let it fail and be told"] B -->|"no"| C{"Is the error transient?"} C -->|"503, 429, network timeout"| C1["Retry On Fail on.
Max Tries 3 to 5"] C -->|"401, 404, a wrong expression"| C2["Retry will not help.
Fix the cause instead"] C1 --> D["For 429, push Wait Between Tries
to 5000 ms or more"]
There is a companion setting worth knowing about while you are in the same tab. Alongside Retry On Fail, a node's Settings has an On Error option — older versions called it Continue On Fail, so recognize the shape rather than the word. It answers a different question: when this node fails for good, does the whole workflow die? Three choices:
- Stop Workflow — the default. The workflow dies, and the error workflow gets called.
- Continue — the node fails but emits empty data, and the run carries on.
- Continue (using error output) — the node grows an extra error branch, and the failed data goes down it.
Continue suits the case where one node failing is acceptable and you want to keep going: batch-processing 100 customers, 3 of them fail, and the other 97 still have to be finished. It combines with retry — three retries, still failing, then Continue moves on to the next batch.
Pausing on purpose
The Wait node does one very simple thing: it stops the workflow here and only moves on when the time is up. You use it to slow things down, to line up with a clock, or to wait for an outside system to answer. These are the three modes you will use most:
| Resume mode | When it resumes | Typical scenario |
|---|---|---|
| After Time Interval | Waits a number of seconds, minutes, hours or days | 1 second between batches so you do not call too fast; wait 10 minutes after an alert, then check the status |
| At Specified Time | Waits until a specific date and time | “Carry on at 9:00 tomorrow morning”; “do not send the invoice until the 1st of next month” |
| On Webhook Call | Waits until an outside system calls the webhook back | Send an approval link and carry on only after they click it; start a payment and confirm only after their callback |
Wait node also has an On Form Submitted mode, where the workflow stops until someone submits a form that n8n generated — approval and data-collection cases. It is advanced; learn it once you are comfortable with webhooks.The most common pattern by far is Wait paired with Split In Batches to stay under a rate limit. Say you have to call an outside API for 500 records. Firing all 500 at once gets you rate-limited.
flowchart LR A["500 records arrive"] --> B["Split In Batches
batch size 10"] B --> C["HTTP Request
10 calls go out"] C --> D["Wait
After Time Interval, 1 second"] D --> B B --> E["All batches done"]
That comes out at roughly 10 requests a second, which usually sits inside what the other side allows.
A workflow sitting at a Wait node is not a taxi with the engine running and the meter ticking. It is a note left on the counter saying “ring this number at nine tomorrow.” The note costs nothing to leave lying there, whether that is for ten seconds or for a week.
Wait node maxes out at 65 seconds. That is wrong. What the docs actually say is that for “wait times less than 65 seconds, the workflow doesn't offload execution data to the database” — so 65 seconds is the threshold where the mechanism switches, not a maximum. Under 65 seconds the workflow is held in memory and nothing is written to the database. At 65 seconds or more, n8n serializes the execution into the database and wakes it when the time comes. Waits of hours or days work the same way, and At Specified Time can wait until next month.Reading the record after the alert
The error was caught and the alert went out. Next you open the Executions page and find out which node failed and why. This is where most of your debugging happens. There are two ways in: Executions in the main sidebar shows the records for every workflow on the instance, while the Executions tab at the top of an open workflow shows only that one.
Each record carries a colored status, and the colors are the fastest thing on the page to read.
| Color | What it means |
|---|---|
| Green Success | Every node ran through, no errors |
| Red Error | A node died partway and the workflow did not finish |
| Orange or yellow Running | Running right now — a Wait node counting down, for example |
| Gray Waiting | Waiting for a webhook, or for At Specified Time — some external condition |
Click into any red execution and you get a canvas-style view of that run. The failed node is highlighted with a red border, and clicking it shows the input, the output and the error message from that specific run. This is the fastest way to find a bug in n8n.
stateDiagram-v2 state "Error Handler runs" as Handler [*] --> Running: an automatic trigger fires Running --> Waiting: the run reaches a Wait node Waiting --> Running: the time is up, or the callback arrives Running --> Success: every node ran through Running --> Error: a node failed for good Error --> Handler: n8n calls the bound error workflow Success --> [*] Handler --> [*]
Pin data, so testing does not burn your quota
Rerunning a workflow after every small change keeps hitting the real APIs — Gmail, Slack, Sheets. That is slow, and it eats quota you may be paying for. Pin data lets you pin one successful output of a node. On later reruns the workflow takes the pinned data at that node instead of calling the API for real.
Get one successful run, click that node's output, and press the pin icon above it. The output is now fixed. When you want to see what a downstream change did, rerun the workflow and the pinned node hands back the pinned data. Unpin it once you are done building, before you go live — a pinned node in a production workflow is a node that has quietly stopped talking to the outside world.
EXECUTIONS_DATA_MAX_AGE environment variable — 336 hours, or 14 days, by default. Ask whoever runs your instance to raise it if you want longer. Or have Error Handler write the errors into Google Sheets or a database at the same time as it posts to Slack, and you keep them for good.The failure you meet, and the fix that goes with it
Here are the errors that come up most often, matched to what to reach for, so that next time a message appears you can look the fix up rather than work it out.
| Failure scenario | What to use | Approach |
|---|---|---|
| 429 Too Many Requests — calling too fast | Wait plus retry | A Wait node to slow the loop down; turn Retry On Fail on for that node and push Wait Between Tries to 5000 ms or more |
| 503 Service Unavailable, or a brief network error | Retry | Max Tries 3–5, Wait Between Tries 2000 ms |
| 401 Unauthorized — the credential expired | Error workflow | Retry is useless here. Let the error workflow tell you, then go to the Credentials page and re-authorize |
| Network timeout | Retry plus Wait | Retry on fail with a 5s wait, so you do not call again before the other side is back |
Bad data format — an expression returns undefined | A Set or IF node upstream | Retry is useless; it is a logic error. Give a default value with Set, or filter with IF, as the data comes in |
| Sub-workflow not found | An error workflow alert | The sub-workflow was deleted or its ID changed. The alert tells you to fix the reference in the Execute Workflow node |
| The Webhook Trigger never arrives | An error workflow cannot catch this | A webhook nobody calls raises no error at all — nothing failed. Add a Schedule workflow that checks every hour how many entries the last hour should have had |
| The Gmail Trigger polls too slowly | Not an error, a design issue | Change Poll Times to Every Minute. If you truly need it instant, find a service with a push webhook |
The other category is “I set it up and nothing happened.” Work through these and you can usually rescue it yourself.
| Symptom | Likely cause | How to fix it |
|---|---|---|
| The workflow died but the error workflow never fired | That workflow's Settings has no Error Workflow set, or the Error Handler workflow is not Active | Go back to the failing workflow → Settings → pick Error Handler in the Error Workflow dropdown. Then open Error Handler and check that Active in the top right is green |
| Five retries and every one of them failed | The error was never transient — an expired credential, data that was wrong to begin with, a logic error | Retry only postpones the failure; it solves nothing at the root. Read the error message in Executions. If it is 401, 400, 404 or a JavaScript error, stop leaning on retry and fix the cause |
The workflow sits at the Wait node and never moves |
On Webhook Call mode but the other side never called the URL; a wait over 65 seconds on SQLite lost in a restart; or At Specified Time with the wrong timezone | Check the webhook URL you gave them is correct and that they really fire the callback; move long waits to an external Postgres database; check At Specified Time against Settings → Timezone |
| The Executions page does not show the run you just did | The page was not refreshed; the workflow was only just made Active and the schedule has not come round; or the retention window has passed | Press F5; check whether the schedule's next time has arrived; have an admin check whether EXECUTIONS_DATA_MAX_AGE was shortened |
Every value in the Slack message is undefined |
The expression has the wrong field path | Run the Error Trigger once for real — trigger it with a workflow that genuinely fails — then look at what $json actually contains in the output panel and copy the field names from there |
| The error workflow died too, because the Slack node failed | Error Handler has no layer above it. It never recurses, by design | Keep Error Handler simple with few dependencies, and stay away from APIs that fail easily. Where it matters, add a second alert inside Error Handler — email as a backup when Slack fails |
| You turned retry on and three copies of the email went out | That node is a non-idempotent operation, so a retry repeats the side effect | Turn retry off on nodes that send mail or create records, and put a “check whether it exists first” node in front instead. Or let that node's failure fail the whole workflow, and have the error workflow tell you to handle it by hand |
| The Error Workflow dropdown is empty | Your Error Handler workflow has no Error Trigger node in it |
Go back to Error Handler and check the first node is an Error Trigger — not Manual, not Webhook. n8n only lists workflows containing an Error Trigger in that dropdown |
Do I have to have an error workflow?
Can retry cause repeated execution and side effects, like two emails or two charges?
Does the Wait node use resources? Can I wait a whole day?
Can several workflows share one error workflow?
{{ $json.workflow.name }} in the message shows which one died. The one exception: if one class of workflow has to alert different people — billing workflows tell finance, operations workflows tell customer support — build two or three error workflows and bind each to its own audience.What happens if the error workflow itself dies?
What is Continue On Fail, and how does it differ from retry?
Does an error workflow only fire for Active workflows? Does a failure from my own Execute Workflow count?
Execute Workflow never starts the Error Trigger. This is not an option, it is a hard rule. So to test an error workflow you have to set the parent workflow to Active, let a Schedule, Webhook or Event Trigger start it automatically, and make it genuinely fail. The design is deliberate — it keeps Slack from being carpet-bombed while you build — but it does mean you cannot test an error workflow by pressing a button on the canvas.How do I use the Wait node's On Webhook Call mode?
Wait node produces a URL — one for Test, one for Production — and the workflow stops there until that URL is called. The classic case: the workflow emails an approval link, the button in the email points at the Wait node's URL, your manager clicks it, the Wait node receives the webhook and the workflow carries on into the approved logic. This is advanced; After Time Interval is enough for most people to start with. Never paste a real production webhook URL into a document or a chat — anyone holding it can resume that execution.What is left to do once error handling is set up?
Where to go from here
Your workflows can now be left running.
That is flow control finished: branching, merging, and the net underneath. Part 13 turns to the data itself — shaping what comes out of one node into what the next one expects, with the Set node.
Part 12 of the Woow n8n Onboarding Guide series on the Apporo blog.
Adapted from the Woow n8n Onboarding Guide, produced by WoowTech and released under CC BY 4.0. This adaptation is published by Apporo under the same licence.
Light · Air · Water · Control · apporo