Skip to Content

Back up the right things, then find out what is actually slow

the last chapter is the safety chapter
Node-RED Guide · Part 12

Back up the right things, then find out what is actually slow

Eleven parts built flows that talk to Home Assistant, call an HTTP API, and publish to MQTT. This part is about the day one of those stops working — because a credential got rotated, a restore did not go the way you expected, or the editor just feels slow and you do not know which of five layers to blame. Everything below is pinned to one exact configuration: Add-on 22.0.1, its bundled Node-RED 5.0.2, and HA WebSocket nodes 0.80.3. Where a behavior depends on that pinning, this part says so, because the same setting can mean something different two releases from now.

3 keys
credential_secret, a Project secret, and the TLS private key under /ssl — three different secrets, and none of them belongs stored next to the data it protects
0.80.3
The pinned HA WebSocket node version. Its own prerequisites — HA 2024.3+, Node-RED 3.1.1+ — are stricter than the Add-on manifest's 2023.3.0 floor
0 = unlimited
The default for nodeMaxMessageBufferLength. A setting you never touched is not a limit you can point to during an incident
Why this is a chapter, not a checkbox

A backup is not a copy of flows.json, and security is not a login screen

This Add-on holds more than your flows. It holds encrypted credentials, it can reach the Home Assistant API and Supervisor, and depending on how you set it up it can also touch several mapped host directories and the host network directly. A leak, an accidental deletion, or a mismatched setting at any one of those layers can stop your automations or hand a capability to a node that should never have had it. Copying flows.json answers none of that on its own.

The two chapters behind this part — backup and security in one, performance and troubleshooting in the other — both work from the same discipline: name the layer, name the failure that layer actually produces, and name the smallest control that closes it. Five layers cover the backup and security side.

LayerPrimary assetsTypical failureMinimum control
DataFlows, credentials, settings, contextLoss, corruption, or an overwrite with the wrong versionEncrypted backups, a retention schedule, restore drills
Keyscredential_secret, the Project secret, the TLS private keyExposed next to the data it protects, or lost so nothing can decryptSeparate storage, no arbitrary rotation, access records
VersionsAdd-on, Node-RED, node, and package versionsUnknown nodes or schema mismatches after a restoreA pinned version inventory and recorded package sources
ExposureIngress, the direct port, HTTP, Dashboard, MQTTTreating editor authentication as authentication for everythingChecking authentication, TLS, and source limits route by route
PermissionsSupervisor, HA API, host network, UART, mapped directoriesA third-party node reaching further than its job needsThe fewest nodes, the fewest destinations, the fewest writable paths
In plain terms

A backup that only copies flows.json is like photocopying the front page of a contract and calling it the whole agreement. The page you kept looks complete. It just does not have the signatures, the appendix, or the key that made any of it enforceable in the first place.

Read this after, not instead of. This part assumes you already have a working Add-on install (Part 1), understand flow architecture and reuse, and know the safe testing habits — Inject, Debug, Catch — from Part 9's debugging chapter. If you cannot yet tell which nodes in your own flows cause a real-world side effect, that is the thing to fix before you run any restore drill below.
Two secrets, not one

credential_secret decides whether your backup is useful, or just ciphertext

Node-RED encrypts every credential you type into a config node — a password, a token, an API key — before it ever touches disk. The Add-on maps its credential_secret option straight onto Node-RED's own credentialSecret setting. The official documentation is blunt about it: store this value securely, and once it is set, do not change it on a whim. Change it and every credential already encrypted with the old value stops decrypting. Not stolen. Not corrupted. Just unreadable, forever, unless you re-enter every password by hand.

In plain terms

Rotating credential_secret without a plan is not like changing a website password. It is like changing the lock on a filing cabinet and then losing the old key before you have moved anything out. The cabinet is not more secure. It is just closed, permanently, on whatever was inside it.

What belongs in your recovery notes is a reference, never the value:

Add-on credential secret reference = PLACEHOLDER_SECRET_RECORD_ID
Project credential secret reference = PLACEHOLDER_PROJECT_SECRET_RECORD_ID
Backup reference = PLACEHOLDER_BACKUP_ID

Those are record IDs pointing into a secrets manager you already trust — not the secrets themselves. Never write the actual value into a flow, a chat message, a support ticket, or a note titled “backup info.”

Projects use a completely different key. If you manually enable Node-RED's Projects feature, credential_secret stays a required Add-on option — but Node-RED quietly ignores it. Project credentials are decrypted with that Project's own secret instead. The two are not interchangeable, and neither one substitutes for the other during a restore.
flowchart TD
  A["Is Projects enabled
on this install?"] -->|"no, the default"| B["The credentials file next to
flows.json is decrypted by
credential_secret"] A -->|"yes, turned on by hand"| C["The Project credential file is
decrypted by the Project's
own secret instead"] B --> D["credential_secret stays a
required Add-on option
either way"] C --> D
One question, one answer either wayThe diamond at the top is the only decision in this picture, and both of its answers flow down into the same last box. That box is the trap: credential_secret is still there in your Add-on options even on the Projects branch, where it does nothing at all — the box you actually need on that side is the one above it, the Project's own secret.

Sanitize before you export, never trust the credentials panel alone

A flow export, a Library entry, or a Subflow can carry server references, environment identifiers, or node properties you did not mean to share — check Subflow instance properties and parent flow context too, not just the top-level flow. Before you post an export anywhere, replace every real value with a complete placeholder: YOUR_HA_SERVER, YOUR_ENTITY_ID, YOUR_DEVICE_ID, YOUR_AREA_ID, YOUR_MQTT_BROKER, YOUR_MQTT_TOPIC. Strip credential blocks, Authorization headers, cookies, webhook IDs, zone and location data, file paths, and any captured Debug output while you are at it.

A few node types are worth a second look before they go into a shared example: Render Template, API, Webhook, Zone, the deprecated Entity node, Server, Device Config, Entity Config, and Update Config. Update Config in particular changes a companion entity's own configuration or state metadata — keep it disabled in anything you post, and leave only placeholders where the real values were.

What a backup has to contain

Six things live under /config, and a drill that proves you can get them back

The Add-on wrapper fixes a handful of paths: userDir=/config/, nodesDir=/config/nodes, flowFile=flows.json. The backend listens only on 127.0.0.1:46836, and the flow HTTP root is fixed at /endpoint. So “the configuration directory” means, at minimum, /config/flows.json, the credentials file paired with it, /config/settings.js, /config/package.json, /config/nodes/, and whatever context storage you have configured. That credentials filename is derived from the flow filename — do not go hunting for it by guessing a name, and do not select backup files by pattern-matching one.

addon_config:rw is what maps that persistent data to /config. homeassistant_config:rw, media:rw, and share:rw are additional capabilities the Add-on can hold — they do not mean Node-RED actually writes to any of those locations, only that it could if a flow told it to. The ssl mount supplies files for direct TLS; if it holds a private key, that key must never be copied into a flow, a log, or a shared directory.

ItemFixed location or sourceBackup and restore considerations
Flows/config/flows.jsonCompare it offline first, and keep every side-effecting node disabled once it is restored
CredentialsThe credentials file paired with the flow fileYou need both the ciphertext and the matching secret — never paste either into a support ticket
Runtime settings/config/settings.jsThe wrapper overwrites specific backend, path, authentication, and TLS keys; do not paste in upstream defaults
Node declarations/config/package.json, the Add-on optionsnode_modules is not in the backup — rebuild it from a pinned, trusted inventory
Local nodes/config/nodesReview the source. Local modules run with the same permissions as the Node-RED process itself
ContextMemory, or a configured localfilesystem storeMemory does not survive a restart, and even a disk store is not guaranteed durable on every assignment
External stateHome Assistant, MQTT, InfluxDB, devicesNone of it is part of a Node-RED backup. A restore must not replay commands or assume the outside world matches

The manifest also excludes node_modules from backups on purpose, with backup_exclude. Keep package.json and a pinned name-and-version list for every custom npm_packages and system_packages entry — that list, not the backup file, is what tells you the backup does not secretly contain the whole dependency tree.

Prove the backup works before you need it to

  1. Step 1

    Write a read-only inventory

    Add-on 22.0.1, Node-RED 5.0.2, HA WebSocket 0.80.3, every additional package's pinned version, whether Projects is enabled, the context store in use, which mapped directories your flows actually touch. Leave out tokens, passwords, private keys, real URLs, and any entity, device, or area ID.

  2. Step 2

    Keep the data and its keys apart

    Back up the Add-on data with an approved Home Assistant backup mechanism. Store credential_secret, and the Project secret if you use one, in an approved secrets manager — on a different host than the backup itself, and never committed to Git.

  3. Step 3

    Restore into an environment that cannot phone home

    No connection to production Home Assistant, MQTT, HTTP, TCP, UDP, WebSocket, a database, or UART. Every Action, API, Fire Event, Update Config, file-output, email, Cast, Modbus, serial, and network node stays disabled, and so does every scheduled trigger.

  4. Step 4

    Validate with read-only checks only

    Compare flow counts, node types, and package versions against your inventory. Confirm the credentials actually decrypt. Confirm the settings and context stores exist. A manual Inject feeding a pure data-processing chain, with Debug limited to a few fields, is as far as this step goes — nothing that writes or calls out.

  5. Step 5

    Write down what you found, then schedule the real thing separately

    Record the backup time, the restored version, anything missing, who verified it, and the rollback method — no secret values. A production restore is its own change window, only after the isolated drill above has actually passed.

flowchart TD
  A["Restore into an isolated environment"] --> B{"Is every side-effect node
disabled or disconnected?"} B -->|"no"| C["Disable Action, API, Fire Event,
MQTT Out, HTTP Request, file,
email, Modbus, serial nodes"] C --> B B -->|"yes"| D{"Do flow counts, node types, and
package versions match
your inventory?"} D -->|"no"| E["Investigate the mismatch.
Do not deploy"] D -->|"yes"| F{"Do the credentials decrypt with
the secret you recorded?"} F -->|"no"| G["This recovery point fails.
Reissue the external credentials"] F -->|"yes"| H["Read-only validation passed.
A production restore is a separate,
approved change window"]
Three ways this can end, and only one is a passThe loop at the top is the only edge that comes back on itself — leaving a side-effect node enabled sends you back to the same question, not forward. From there the drill only ever lands in one of three boxes at the bottom, and the middle one, a decryption failure, is the one that means going back to the secrets manager rather than trying another restore.
Do not deploy right after an incident. Isolate first, and preserve the timeline and de-identified logs before you touch anything. If a credential, a Project secret, an HA token, MQTT credentials, HTTP Basic Auth, a TLS private key, or a webhook ID may have leaked, rotate in this order: the external system's credential first, then the Node-RED side, then validate with flows still disabled. Replacing credential_secret itself is not routine rotation — it breaks decryption of everything already encrypted with the old value.

Projects and Git are one layer, not a replacement for the others

The Add-on's settings.js sets editorTheme.projects.enabled=false by default. Turn Projects on and you get a Git-backed workflow — but a Git commit, a reachable remote, and a complete Add-on backup are three different things. Projects splits into three separate assets: the plaintext credentials themselves; the Project credential file, encrypted with the Project secret and the one piece that is safe to track in Git; and the Project secret, which never enters the repository at all.

  • A private repository is still not a secrets manager. Minimize who and what can reach it — members, deploy keys, automation tokens — and never commit plaintext credentials, the Project secret, TLS keys, backup files, or real endpoint and environment IDs.
  • Keep the Project secret somewhere else entirely. The Add-on's credential_secret does not stand in for it, and it never travels alongside the encrypted Project credential file in Git. Confirm you actually have the matching secret before restoring a Project — do not gamble on “resetting the key” instead.
  • If your policy will not allow even ciphertext in Git, then Git cannot restore your Project credentials at all, and you need a separately backed-up, separately verified credential file.
  • If the remote goes missing, preserve a read-only copy of the local Project directory first. Do not force-push, delete .git, or reinitialize the Project over its own history — that is a job for an authorized maintainer once the identity and target remote are both confirmed.
One login does not cover the rest

Ingress protects one door. It says nothing about the other six

Assuming that logging into the editor secures the whole Add-on is the single most common mistake in this chapter, and it is worth walking through why: each surface below is a different door, checked by a different mechanism, and some of them are not checked at all unless you set that up yourself.

SurfaceBehavior in 22.0.1Required controls
Ingressingress=true, dynamic ingress_port=0, ingress_stream=true — NGINX accepts only the Supervisor Ingress sourceApply least privilege to HA accounts. Do not assume the direct endpoint gets the same protection
Direct editorContainer port 80/tcp can map to host port 1880; / uses Supervisor authentication by defaultMap the port only when you need it, and leave leave_front_door_open false or unset
TLSssl defaults to true and only affects the direct listener; certfile and keyfile are read from /sslConfirm the pair matches and is current. This does not issue or renew certificates, and it does not touch Ingress
Flow HTTP/endpoint/ does not use the direct editor's Supervisor authentication; http_node supplies Basic AuthSet your own strong credentials, TLS, and input limits. Never mistake this for editor authentication
Static contenthttp_static protects static content onlyDo not assume that protection reaches HTTP nodes, the editor, WebSockets, or any other route
DashboardFlowFuse Dashboard 2 is optional, not bundled, and served by Node-RED itself once installedVerify HTTP and WebSocket authentication route by route — compatibility metadata is not a test
MQTT and other networksMQTT, HTTP, WebSocket, TCP, UDP, and TLS config nodes can all produce external I/OBroker ACLs, a dedicated client, TLS verification, least-privilege topic access
flowchart TD
  N["Node-RED Add-on 22.0.1"] --> ING["Ingress: NGINX accepts only
the Supervisor Ingress source"] N --> DIR["Direct editor, port 80/tcp,
mapped to host 1880,
Supervisor auth by default"] N --> EP["Flow HTTP root /endpoint,
guarded only by http_node's
own Basic Auth, if you set it"] N --> ST["Static content under
http_static, files only"] N --> DASH["Optional FlowFuse Dashboard 2,
routes served by Node-RED itself"]
Five doors off the same boxFive boxes hang directly off the add-on, and every one of them is a terminal box — nothing here feeds into anything else. /endpoint is the one to read twice: it is the only door in this picture with no Supervisor authentication behind it at all, so whatever http_node's own Basic Auth is set to is the entire lock.
In plain terms

Think of the shop's front door, the loading dock round the back, and a vending machine bolted to the outside wall. The Supervisor guards the front door and checks every badge. Nobody checks the vending machine unless you personally fit a lock to it — that vending machine is /endpoint, and http_node's Basic Auth is the only lock it will ever have.

What the manifest grants is a ceiling, not a to-do list

hassio_api=true with hassio_role=manager hands the Add-on powerful Supervisor capabilities. homeassistant_api=true reaches the HA Core API. auth_api=true lets the wrapper use Supervisor authentication. host_network=true widens what the Add-on can reach on your local network, and uart=true opens serial hardware access. None of that is a requirement to use all of it — it is the ceiling a misbehaving or malicious node could reach. Never write a Supervisor or HA token into a flow or into Debug output, and keep the list of additional nodes, outbound destinations, and writable directories as short as the job actually needs.

Flows can write to homeassistant_config:rw, media:rw, and share:rw. Allowlist the exact filenames and paths you use, to rule out path traversal, accidental overwrites, and a secret ending up in a shared directory by mistake. UART and Modbus writes can move something in the physical world — validate them in an isolated environment with no hardware attached, every time.

Packages and init_commands run inside the trust boundary

At startup the wrapper installs system_packages, then runs npm install --omit=dev --omit=optional for npm_packages, and finally passes each line of init_commands to eval. A failure at any of those three stages aborts the whole startup. Because init_commands runs at every restart, not just the first one, keep it empty unless every command has gone through change review, is repeatable, and contains no secret.

A few categories of pinned package are worth flagging by what they do, whatever you name them: one hashes HTTP Basic Auth passwords for the wrapper — a hash is not a publishable password. Some can control devices or send outbound messages directly, and belong disabled during any restore drill. Others manage database credentials, query scope, and retention on their own. A few touch OT or serial hardware, where the risk compounds with uart=true. And a couple only read untrusted content or probe the network, where the control that matters is limiting sources, sizes, and destinations — not the package name.

Which layer is actually slow

“It's slow” is five different problems wearing the same word

“Node-RED is slow” could mean the editor is drowning in Debug output, a message got too large somewhere, a sequence buffer keeps growing, external I/O is stuck waiting, or V8 is garbage-collecting old space too often. It does not automatically mean you need more memory. The fix is the same one this whole part keeps returning to: find the layer before you touch a setting.

LayerExample symptomCheck firstDo not do first
Supervisor / Add-onA startup loop, or initialization that abortsStartup time, the first error, recent option changesClear settings, or hide the first error under repeated restarts
Node.js / runtimeFrequent GC, event-loop latency, or the process exitingMemory trend, message rate, the old-space settingSet the heap size close to the host's total RAM
Flow / nodeDuplicate messages, buffer growth, an unknown nodeNode status, Catch output, message shape, Deploy timeDelete an unknown node, or run a Full Deploy to see what happens
External I/OTimeouts, reconnects, delayed dataDestination status, request duration, retry rate, queue depthRun unauthorized probes, or disable TLS verification to test faster
Reverse proxy / authIngress opens, but one endpoint failsIngress vs. direct access, port, path, TLS, and auth, checked separatelyTurn on leave_front_door_open
flowchart TD
  A["Something feels slow or is failing"] --> B{"Did it start right after
an option change or upgrade?"} B -->|"yes"| C["Supervisor / Add-on layer:
first error, recent option changes"] B -->|"no"| D{"Is memory or CPU
climbing over time?"} D -->|"yes"| E["Node.js / runtime layer:
GC frequency, message rate,
the old-space setting"] D -->|"no"| F{"Is it one specific
flow or node?"} F -->|"yes"| G["Flow / node layer:
status, Catch output,
buffer depth"] F -->|"no"| H["External I/O, or the reverse
proxy and auth layer:
check both separately"]
Four questions, four leavesEvery diamond has exactly two ways out, and every one of those routes ends in a leaf box with nothing after it — four leaves in total. The last leaf covers two of the five layers from the table above at once, which is deliberate: if the first three questions all came back “no,” you have not yet separated an I/O problem from an authentication one, and that split still has to happen inside that box.

Debug output and message cloning

A Debug node should show one sanitized field, not the whole message object. The Add-on's settings template sets debugMaxLength to 1000 characters — but that only truncates the display. It proves nothing about whether the full object upstream was ever cloned or retained. Node-RED can clone a message every time it branches, and large nested objects, Buffers, and many branches all make that more expensive. Images, audio, large JSON values, and a live msg.req or msg.res do not belong in routine Debug output or in context, ever.

RBE: a stateful filter, not a stateless conversion

RBE — shown as Filter in the palette — passes a message through only when a value changes, or when a numeric deadband or narrowband rule allows it. It keeps the previous comparison value per topic, which makes it stateful, not a pass-through. You choose the property to compare, typically msg.payload, and the topic property. A deadband comparison expects a number, and because the node reaches for parseFloat under the hood, reject strings and objects before they get anywhere near it rather than trusting permissive parsing to do the right thing.

msg.payload — the usual property RBE compares against its stored previous value.
msg.reset — forces RBE to forget its stored value for that topic and treat the next message as new.

Test it against an unchanged value, a value right at the threshold, a value past it, a different topic entirely, a msg.reset, and what happens to the comparison state after a restart or a Deploy. That last case matters more than it looks: whether the very next message gets through can depend on state you never explicitly set, so nothing downstream should assume a particular prior state that was never actually tested.

Delay, Trigger, Join — give every wait a bound

Delay — the rate limit and queue policy have to handle your peak load, not your average. Decide up front whether intermediate values are discarded or kept; that is a product decision, not a default to accept.
Trigger — each topic or stream can hold its own timer. Unbounded input keys mean an unbounded number of pending timers. Set a real timeout and a key allowlist.
Split / Join / Batch / Sort — keep msg.parts correct, and cap the number of waiting sequences, items per group, and wait time. nodeMaxMessageBufferLength is the Add-on's own limit for this, and its default of 0 means unlimited — not evidence that a limit exists.
HA Wait Until — give every wait an explicit timeout and a timeout branch. A pile of simultaneous Time triggers is its own performance problem.

Context and external I/O

Node, flow, and global context can live in memory or in a named localfilesystem store. Neither is the place for unbounded arrays, whole event objects, or binary payloads, and a disk-backed store is not guaranteed to write durably on every single assignment — decide the limits on keys, size, retention, and write frequency before you rely on it. HTTP, MQTT, WebSocket, TCP, UDP, serial, files, databases, and HA's Get History or Poll State are all external I/O in this sense: give each one a timeout, a result-size limit, a retry limit, and a backoff policy. Limit the time range and result count on Get History specifically, and reach for an event subscription instead of Poll State wherever the two would do the same job.

Safe Mode and memory limits

Safe Mode is neutral gear, not the engine off

The Add-on option safe_mode: true adds the --safe flag at startup. Node-RED still starts — it just suppresses the initial startup of your flows, giving you a window to fix something before it runs again. What it does not do is stop a Deploy from starting the flows in that deployment, even while safe_mode is still set to true. It is not Home Assistant's own Safe Mode, it does not stop the Add-on, and it changes nothing about editor or admin authentication — Ingress and direct-access rules still apply exactly as before.

In plain terms

Safe Mode is putting the car in neutral, not switching the engine off. The engine keeps running. Nothing moves until you press a pedal — and Deploy is the pedal. Press it, and whatever is in that deployment goes, safe_mode or not.

sequenceDiagram
  participant Op as Operator
  participant AO as Add-on, safe_mode true
  participant NR as Node-RED runtime
  Op->>AO: restart with safe_mode true
  AO->>NR: start with --safe
  NR->>NR: suppress the initial startup of flows
  Op->>NR: click Deploy
  NR->>NR: start every flow in that deployment
Safe Mode ends at the third arrow, not the last oneThe first three arrows are what safe_mode actually buys you: flows stay stopped through the restart. The last two arrows happen regardless — a Deploy from the operator still tells the runtime to start what is in it, which is why disabling side-effect nodes has to happen before that click, not after.

Before the first Deploy after any restore or repair, disable every node with a network call, an HA action, a write, or a schedule attached — or disconnect the wire leading into it. Do not Deploy first and clean up second; a Deploy that already ran cannot be undone by a check that comes after it.

Deploy scope is not a performance switch

Node-RED 5.0.2's editor offers three Deploy scopes: Full, Modified Flows, and Modified Nodes. Which one you pick decides which flows or nodes actually restart, and that in turn affects timers, open connections, context lifecycles, and any message mid-flight. Understand what depends on the thing you are changing before you choose the narrowest scope — if a shared config node's blast radius is not clear to you, do not reach for Modified Nodes just because it looks faster.

max_old_space_size limits one drawer, not the whole cabinet

This Add-on option is an integer in MB. When it is set, the startup script exports exactly one line:

NODE_OPTIONS=--max_old_space_size=PLACEHOLDER_MB
In plain terms

max_old_space_size is turning up the size of one drawer in a filing cabinet, not renting a bigger office. Old space is one section of the V8 heap. The young generation, native add-ons, Buffers, and general overhead all live in other drawers this setting never touches — and the cabinet still has to fit in the room, which is the Home Assistant host's actual RAM.

  1. Record process and container memory, message rate, latency, restarts, and queue depth first — before changing anything.
  2. Fix unbounded Debug output, oversized messages, unbounded context, and unbounded Delay/Trigger/Join buffers before you even consider this setting. More old space does not fix a leak.
  3. If a real need remains, pick a conservative value, change only that, and repeat the same load in an isolated environment.
  4. Watch GC behavior, latency, and host headroom. No improvement means roll back — not raise the number again.
There is no universal value. The right number depends on the host, the other Add-ons sharing it, your flows, and your actual load. Any “best” MB figure offered without measurements behind it is a guess dressed up as advice.

Log level, and what not to paste into a ticket

log_level accepts trace, debug, info, notice, warning, error, and fatal, and the option can be left unset. The official docs recommend info for normal use; the wrapper maps warning onto Node-RED's own warn. Turning verbosity up can print payloads, URLs, headers, entity or device IDs, MQTT topics, file paths, and stack traces straight into the log — use a higher level only for a defined window, then put it back. When you do share a log, keep versions, timestamps, node types, error categories, and a correlation placeholder, and strip everything else: tokens, passwords, Authorization and Cookie headers, encrypted credentials, private keys, webhook IDs, location data, real hostnames and IPs, full payloads, and anything visible in a screenshot's edges.

Upgrading without losing the thread

Six startup stages, and each one has its own way to fail

Before you upgrade anything: finish an Add-on backup and an isolated restore drill, with credential_secret or the Project secret stored separately from both. Export a sanitized flow inventory — Add-on, Node-RED, and HA WebSocket versions, every extra npm and system package, and the optional, not-bundled FlowFuse Dashboard 2 if you use it. Its own metadata calls for Node 14 or newer and Node-RED 3.0.0 or newer, and that compatibility line is not a substitute for testing it yourself in your own environment. Read the target release's notes and breaking changes. Disable HA, network, write, and scheduling nodes, and decide your stop condition and rollback version before you start — not while you are already mid-upgrade.

StageBehavior in 22.0.1Safe response to a failure
Data initializationA fresh /config may migrate data from /homeassistant/node-red; otherwise settings, flows, and node directories are created newInventory both directories before touching anything. Never move files by hand — find the first migration error instead
Theme migrationAn old dark theme is renamed to dark-modernThat is a name change, not flow corruption — do not misdiagnose it as one
Conflicting packagesThe wrapper tries to remove node-red-contrib-home-assistant, node-red-contrib-home-assistant-llat, and node-red-contrib-home-assistant-wsThis is scripted wrapper behavior — do not run your own removal command. If it fails, keep the log and the package inventory
Custom packagesAlpine system_packages install first, then npm npm_packages; any failure aborts startupFind the first package, repository, or architecture error, and roll back only the most recent option change
Custom commandsEach line of init_commands is passed to eval at every startup; any failure aborts startupDisable the most recently added command and return to known-good settings. Never put a secret in a command or a log
Runtime flagssafe_mode becomes --safe; max_old_space_size becomes NODE_OPTIONSConfirm each option exists with the right type, and do not confuse the manifest's init=false with runtime Safe Mode

After an upgrade, stay in Safe Mode and verify the editor, node types, credential decryption, settings, and package loading — all without deploying. Only once every side-effect node is disabled or disconnected does a Deploy of the smallest pure-data flow become the next, separately approved, step. If rollback becomes necessary, it means more than swapping the image back: stop making further changes, preserve sanitized logs from the failed version, and use the Home Assistant Add-on's own restore procedure to land on a matching backup, version, and secret together. Do not feed data a newer version already migrated back into an older runtime, and never hand-edit a node's type or version field inside flow JSON.

The version hedge that actually bites people. HA WebSocket 0.80.3 itself requires HA 2024.3 or newer, Node-RED 3.1.1 or newer, and Node 18.2.0 or newer — a stricter floor than the Add-on manifest's own HA 2023.3.0 installation threshold. Meeting the Add-on's minimum does not automatically mean you meet the node package's minimum. Check both before you assume a “disconnected” HA node is a configuration mistake rather than a version gap.
When it does not add up

The failures that actually happen, matched to what to check first

SymptomLikely causeWhat to do
Credentials don't work after a restoreWrong secret, wrong version, or a Project boundary crossedKeep flows disabled. Verify the backup's version and Project boundary against the secret you have. Don't keep changing credential_secret hoping one attempt works — if no matching secret exists, reissue the external credentials instead
Ingress opens, but /endpoint/ has no authentication you expectedYou're thinking of Supervisor auth, which /endpoint never hadConfirm whether you're on Ingress or the direct port, then test with a read-only, no-real-data request. Check http_node, TLS, and the port mapping — never reach for leave_front_door_open
Unknown nodes appear after a restoreA package version wasn't rebuilt, or doesn't match the inventoryDon't deploy or delete it. Match it against the pinned-package inventory, rebuild the excluded node_modules in isolation, and verify source and version before it loads
The Project remote or its credentials failA mismatched secret, remote identity, or access scopePreserve the local data and Git status as-is. Verify the Project secret and remote identity. Don't force-push, reset the secret, or reinitialize the Project
You suspect a token or key leakedExposure through a log, a backup, Git, or a shared directoryIsolate the Add-on first and preserve a de-identified timeline, then rotate the credential at its source. Deleting the Debug message is not the fix — check backups, Git, and shared directories for the full exposure
Add-on startup stops at a custom package or commandA system_packages, npm_packages, or init_commands failureFind the first error in the log, compare it against recent option changes and the architecture, and roll back only the latest change — not an ad hoc shell command
An HA node reports disconnected or Unauthorized WebSocketThe connection mode setting, or a version below 0.80.3's own floorConfirm “I use the Home Assistant Add-on” is set correctly in the HA Server config. Check HA 2024.3+, Node-RED 3.1.1+, and Node 18.2.0+ — don't copy tokens or send a test Action to debug it
An HTTP node or Dashboard route 404sConfusing Ingress with the direct, separately mapped portHTTP nodes need their own mapped port and live under /endpoint/. Check the port, the TLS pair, and http_node authentication — not TLS verification or leave_front_door_open
A TLS handshake or certificate errorAn expired, mismatched, or wrongly placed fileCheck filename, validity period, subject, and chain under /ssl only. Never expose a private key or disable verification to move faster
Join or Wait Until emits nothing, and memory keeps growingAn unbounded sequence, timeout, or pending-key setSample msg.parts, group size, timeouts, and pending keys. Stop new input and preserve evidence — don't inject more test traffic or clear production context directly
Repeated restarts hid the real first errorOnly later timeouts survived, not the original failureStop restarting. Treat the earliest first error and the last known good time as primary evidence, and label anything after that as secondary. Don't guess the root cause from what's left
Unknown nodes or schema warnings appear after import or upgradeA pinned package version wasn't matched, or a legacy node type has been deprecatedStay in Safe Mode. Don't Deploy, delete nodes, or edit their type manually. Identify the pinned package from the backup inventory and restore a compatible version in an isolated environment first. The legacy HA Entity node is deprecated — handle it only through the migration procedure for the pinned version, and don't create a new one
An upgrade report or escalation is requiredThe issue survives isolated troubleshooting and needs a maintainerStop making repeated changes. Give the Add-on or node-package maintainer pinned versions, the architecture, a reproduction timeline, the first error, whether the problem remains with side effects disabled, a minimal sanitized flow, and the one change already attempted. Never attach tokens, credentials, private keys, real endpoints or IDs, full payloads, browser storage, or unredacted screenshots
Questions people ask

The ones that come up again and again

Can I rotate credential_secret on a schedule, like an access token?
No. The official Add-on documentation is explicit that changing it arbitrarily breaks decryption of every credential already stored under the old value. During an incident, rotate the external system's credential first; if the key itself needs to change, treat it as a controlled migration, not routine hygiene.
Can a private Git repository for Projects replace an Add-on backup?
No. Git may not cover Add-on options, non-Project data, context, or every credential and wrapper state — and a private repository is not a secrets manager on its own. Keep both as separate recovery layers, each verified on its own.
Does logging into Ingress protect HTTP, Dashboard, and MQTT too?
No. Ingress, the direct editor, /endpoint/, static content, Dashboard routes, WebSockets, and MQTT are separate boundaries, each with its own authentication story. Set each one up on its own terms rather than assuming one login covers the rest.
How do I run a restore drill without triggering real smart-home actions?
Restore into an environment with no production network, HA, MQTT, or UART connection. Disable every external I/O node, write, schedule, and executable HA node before the first Deploy. Limit yourself to read-only checks: file integrity, node types, whether credentials decrypt, and pure data-processing chains with a manual Inject.
Can I set max_old_space_size to my host's total RAM?
No. It bounds only V8's old space, not the whole Node.js process, the container, or anything else sharing the host. Fix unbounded messages, buffers, and context first, then test a conservative value under real load with real headroom left for the host.
Does Safe Mode stop the whole Add-on?
No. It starts Node-RED with --safe and only suppresses the initial startup of flows. Any Deploy — even while safe_mode stays true — can still start the flows in that deployment. Disable or disconnect every side-effect path before that first Deploy, not after.
Does truncating Debug output reduce its actual memory cost?
Not necessarily. debugMaxLength limits the display only — it says nothing about whether the full object was cloned or retained upstream. Reduce the message at its source, and separately limit both how often Debug fires and which fields it shows.
Can I just delete and recreate an unknown node after a restore?
Don't. Deleting it loses the original configuration, and a Deploy afterward may start other flows you didn't intend to touch. Stay in Safe Mode, match the node against your pinned-package inventory, and restore a compatible version in isolation, or follow the official migration path if one exists.
My HA node says disconnected right after an upgrade. Where do I look first?
Check the connection mode setting first — the HA Server config needs “I use the Home Assistant Add-on” set correctly. Then check the version floor: HA WebSocket 0.80.3 needs HA 2024.3+, Node-RED 3.1.1+, and Node 18.2.0+, which is stricter than the Add-on manifest's own 2023.3.0 threshold.
Why does the restore drill insist on an isolated environment at all?
Because a restore that reconnects to production before you've verified it can replay commands, fire schedules, or write to real devices before you've confirmed anything is actually correct. Isolation is what turns a restore into a test instead of a second incident.
During an incident, the secrets-manager reference does not match the backup time. Which key should I try first?
Don't guess or overwrite the configuration. Stop the restore and preserve read-only evidence of the backup hash, timestamp, and reference version. Use the secrets manager's own access records to establish the match. If you can't prove a match, mark that recovery point as failed and follow the approved incident-response and external-credential reissuance procedures.
Next

Where to go from here

twelve parts, one working install

You now have a Node-RED install that talks to Home Assistant, and the habits to keep it that way

This closes the series: safe foundations, Home Assistant in practice, and the engineering and operations that keep both running when something goes wrong instead of when everything is fine. If an earlier part sent you back here — a failed restore, a slow flow, an upgrade that would not settle — the fix is almost always in one of the tables above. Keep this part bookmarked longer than any other; it is the one you come back to.

Open the full guide

Part 12 of the WoowTech Complete Node-RED Guide series on the Apporo blog.

Adapted from the WoowTech Complete Node-RED Guide, produced by WoowTech and released under CC BY 4.0. This adaptation is published by Apporo under the same licence.

Light · Air · Water · Control · apporo

MQTT in Node-RED, and the line between what shipped and what you add