Back up the right things, then find out what is actually slow
Eleven parts built flows that talk to Home Assistant, call an HTTP API, and publish to MQTT. This part is about the day one of those stops working — because a credential got rotated, a restore did not go the way you expected, or the editor just feels slow and you do not know which of five layers to blame. Everything below is pinned to one exact configuration: Add-on 22.0.1, its bundled Node-RED 5.0.2, and HA WebSocket nodes 0.80.3. Where a behavior depends on that pinning, this part says so, because the same setting can mean something different two releases from now.
credential_secret, a Project secret, and the TLS private key under /ssl — three different secrets, and none of them belongs stored next to the data it protectsnodeMaxMessageBufferLength. A setting you never touched is not a limit you can point to during an incidentA backup is not a copy of flows.json, and security is not a login screen
This Add-on holds more than your flows. It holds encrypted credentials, it can reach the Home Assistant API and Supervisor, and depending on how you set it up it can also touch several mapped host directories and the host network directly. A leak, an accidental deletion, or a mismatched setting at any one of those layers can stop your automations or hand a capability to a node that should never have had it. Copying flows.json answers none of that on its own.
The two chapters behind this part — backup and security in one, performance and troubleshooting in the other — both work from the same discipline: name the layer, name the failure that layer actually produces, and name the smallest control that closes it. Five layers cover the backup and security side.
| Layer | Primary assets | Typical failure | Minimum control |
|---|---|---|---|
| Data | Flows, credentials, settings, context | Loss, corruption, or an overwrite with the wrong version | Encrypted backups, a retention schedule, restore drills |
| Keys | credential_secret, the Project secret, the TLS private key | Exposed next to the data it protects, or lost so nothing can decrypt | Separate storage, no arbitrary rotation, access records |
| Versions | Add-on, Node-RED, node, and package versions | Unknown nodes or schema mismatches after a restore | A pinned version inventory and recorded package sources |
| Exposure | Ingress, the direct port, HTTP, Dashboard, MQTT | Treating editor authentication as authentication for everything | Checking authentication, TLS, and source limits route by route |
| Permissions | Supervisor, HA API, host network, UART, mapped directories | A third-party node reaching further than its job needs | The fewest nodes, the fewest destinations, the fewest writable paths |
A backup that only copies flows.json is like photocopying the front page of a contract and calling it the whole agreement. The page you kept looks complete. It just does not have the signatures, the appendix, or the key that made any of it enforceable in the first place.
credential_secret decides whether your backup is useful, or just ciphertext
Node-RED encrypts every credential you type into a config node — a password, a token, an API key — before it ever touches disk. The Add-on maps its credential_secret option straight onto Node-RED's own credentialSecret setting. The official documentation is blunt about it: store this value securely, and once it is set, do not change it on a whim. Change it and every credential already encrypted with the old value stops decrypting. Not stolen. Not corrupted. Just unreadable, forever, unless you re-enter every password by hand.
Rotating credential_secret without a plan is not like changing a website password. It is like changing the lock on a filing cabinet and then losing the old key before you have moved anything out. The cabinet is not more secure. It is just closed, permanently, on whatever was inside it.
What belongs in your recovery notes is a reference, never the value:
Add-on credential secret reference = PLACEHOLDER_SECRET_RECORD_IDProject credential secret reference = PLACEHOLDER_PROJECT_SECRET_RECORD_IDBackup reference = PLACEHOLDER_BACKUP_IDThose are record IDs pointing into a secrets manager you already trust — not the secrets themselves. Never write the actual value into a flow, a chat message, a support ticket, or a note titled “backup info.”
credential_secret stays a required Add-on option — but Node-RED quietly ignores it. Project credentials are decrypted with that Project's own secret instead. The two are not interchangeable, and neither one substitutes for the other during a restore.flowchart TD A["Is Projects enabled
on this install?"] -->|"no, the default"| B["The credentials file next to
flows.json is decrypted by
credential_secret"] A -->|"yes, turned on by hand"| C["The Project credential file is
decrypted by the Project's
own secret instead"] B --> D["credential_secret stays a
required Add-on option
either way"] C --> D
credential_secret is still there in your Add-on options even on the Projects branch, where it does nothing at all — the box you actually need on that side is the one above it, the Project's own secret.Sanitize before you export, never trust the credentials panel alone
A flow export, a Library entry, or a Subflow can carry server references, environment identifiers, or node properties you did not mean to share — check Subflow instance properties and parent flow context too, not just the top-level flow. Before you post an export anywhere, replace every real value with a complete placeholder: YOUR_HA_SERVER, YOUR_ENTITY_ID, YOUR_DEVICE_ID, YOUR_AREA_ID, YOUR_MQTT_BROKER, YOUR_MQTT_TOPIC. Strip credential blocks, Authorization headers, cookies, webhook IDs, zone and location data, file paths, and any captured Debug output while you are at it.
A few node types are worth a second look before they go into a shared example: Render Template, API, Webhook, Zone, the deprecated Entity node, Server, Device Config, Entity Config, and Update Config. Update Config in particular changes a companion entity's own configuration or state metadata — keep it disabled in anything you post, and leave only placeholders where the real values were.
Six things live under /config, and a drill that proves you can get them back
The Add-on wrapper fixes a handful of paths: userDir=/config/, nodesDir=/config/nodes, flowFile=flows.json. The backend listens only on 127.0.0.1:46836, and the flow HTTP root is fixed at /endpoint. So “the configuration directory” means, at minimum, /config/flows.json, the credentials file paired with it, /config/settings.js, /config/package.json, /config/nodes/, and whatever context storage you have configured. That credentials filename is derived from the flow filename — do not go hunting for it by guessing a name, and do not select backup files by pattern-matching one.
addon_config:rw is what maps that persistent data to /config. homeassistant_config:rw, media:rw, and share:rw are additional capabilities the Add-on can hold — they do not mean Node-RED actually writes to any of those locations, only that it could if a flow told it to. The ssl mount supplies files for direct TLS; if it holds a private key, that key must never be copied into a flow, a log, or a shared directory.
| Item | Fixed location or source | Backup and restore considerations |
|---|---|---|
| Flows | /config/flows.json | Compare it offline first, and keep every side-effecting node disabled once it is restored |
| Credentials | The credentials file paired with the flow file | You need both the ciphertext and the matching secret — never paste either into a support ticket |
| Runtime settings | /config/settings.js | The wrapper overwrites specific backend, path, authentication, and TLS keys; do not paste in upstream defaults |
| Node declarations | /config/package.json, the Add-on options | node_modules is not in the backup — rebuild it from a pinned, trusted inventory |
| Local nodes | /config/nodes | Review the source. Local modules run with the same permissions as the Node-RED process itself |
| Context | Memory, or a configured localfilesystem store | Memory does not survive a restart, and even a disk store is not guaranteed durable on every assignment |
| External state | Home Assistant, MQTT, InfluxDB, devices | None of it is part of a Node-RED backup. A restore must not replay commands or assume the outside world matches |
The manifest also excludes node_modules from backups on purpose, with backup_exclude. Keep package.json and a pinned name-and-version list for every custom npm_packages and system_packages entry — that list, not the backup file, is what tells you the backup does not secretly contain the whole dependency tree.
Prove the backup works before you need it to
-
Step 1
Write a read-only inventory
Add-on 22.0.1, Node-RED 5.0.2, HA WebSocket 0.80.3, every additional package's pinned version, whether Projects is enabled, the context store in use, which mapped directories your flows actually touch. Leave out tokens, passwords, private keys, real URLs, and any entity, device, or area ID.
-
Step 2
Keep the data and its keys apart
Back up the Add-on data with an approved Home Assistant backup mechanism. Store
credential_secret, and the Project secret if you use one, in an approved secrets manager — on a different host than the backup itself, and never committed to Git. -
Step 3
Restore into an environment that cannot phone home
No connection to production Home Assistant, MQTT, HTTP, TCP, UDP, WebSocket, a database, or UART. Every Action, API, Fire Event, Update Config, file-output, email, Cast, Modbus, serial, and network node stays disabled, and so does every scheduled trigger.
-
Step 4
Validate with read-only checks only
Compare flow counts, node types, and package versions against your inventory. Confirm the credentials actually decrypt. Confirm the settings and context stores exist. A manual Inject feeding a pure data-processing chain, with Debug limited to a few fields, is as far as this step goes — nothing that writes or calls out.
-
Step 5
Write down what you found, then schedule the real thing separately
Record the backup time, the restored version, anything missing, who verified it, and the rollback method — no secret values. A production restore is its own change window, only after the isolated drill above has actually passed.
flowchart TD
A["Restore into an isolated environment"] --> B{"Is every side-effect node
disabled or disconnected?"}
B -->|"no"| C["Disable Action, API, Fire Event,
MQTT Out, HTTP Request, file,
email, Modbus, serial nodes"]
C --> B
B -->|"yes"| D{"Do flow counts, node types, and
package versions match
your inventory?"}
D -->|"no"| E["Investigate the mismatch.
Do not deploy"]
D -->|"yes"| F{"Do the credentials decrypt with
the secret you recorded?"}
F -->|"no"| G["This recovery point fails.
Reissue the external credentials"]
F -->|"yes"| H["Read-only validation passed.
A production restore is a separate,
approved change window"]
credential_secret itself is not routine rotation — it breaks decryption of everything already encrypted with the old value.Projects and Git are one layer, not a replacement for the others
The Add-on's settings.js sets editorTheme.projects.enabled=false by default. Turn Projects on and you get a Git-backed workflow — but a Git commit, a reachable remote, and a complete Add-on backup are three different things. Projects splits into three separate assets: the plaintext credentials themselves; the Project credential file, encrypted with the Project secret and the one piece that is safe to track in Git; and the Project secret, which never enters the repository at all.
- A private repository is still not a secrets manager. Minimize who and what can reach it — members, deploy keys, automation tokens — and never commit plaintext credentials, the Project secret, TLS keys, backup files, or real endpoint and environment IDs.
- Keep the Project secret somewhere else entirely. The Add-on's
credential_secretdoes not stand in for it, and it never travels alongside the encrypted Project credential file in Git. Confirm you actually have the matching secret before restoring a Project — do not gamble on “resetting the key” instead. - If your policy will not allow even ciphertext in Git, then Git cannot restore your Project credentials at all, and you need a separately backed-up, separately verified credential file.
- If the remote goes missing, preserve a read-only copy of the local Project directory first. Do not force-push, delete
.git, or reinitialize the Project over its own history — that is a job for an authorized maintainer once the identity and target remote are both confirmed.
Ingress protects one door. It says nothing about the other six
Assuming that logging into the editor secures the whole Add-on is the single most common mistake in this chapter, and it is worth walking through why: each surface below is a different door, checked by a different mechanism, and some of them are not checked at all unless you set that up yourself.
| Surface | Behavior in 22.0.1 | Required controls |
|---|---|---|
| Ingress | ingress=true, dynamic ingress_port=0, ingress_stream=true — NGINX accepts only the Supervisor Ingress source | Apply least privilege to HA accounts. Do not assume the direct endpoint gets the same protection |
| Direct editor | Container port 80/tcp can map to host port 1880; / uses Supervisor authentication by default | Map the port only when you need it, and leave leave_front_door_open false or unset |
| TLS | ssl defaults to true and only affects the direct listener; certfile and keyfile are read from /ssl | Confirm the pair matches and is current. This does not issue or renew certificates, and it does not touch Ingress |
| Flow HTTP | /endpoint/ does not use the direct editor's Supervisor authentication; http_node supplies Basic Auth | Set your own strong credentials, TLS, and input limits. Never mistake this for editor authentication |
| Static content | http_static protects static content only | Do not assume that protection reaches HTTP nodes, the editor, WebSockets, or any other route |
| Dashboard | FlowFuse Dashboard 2 is optional, not bundled, and served by Node-RED itself once installed | Verify HTTP and WebSocket authentication route by route — compatibility metadata is not a test |
| MQTT and other networks | MQTT, HTTP, WebSocket, TCP, UDP, and TLS config nodes can all produce external I/O | Broker ACLs, a dedicated client, TLS verification, least-privilege topic access |
flowchart TD N["Node-RED Add-on 22.0.1"] --> ING["Ingress: NGINX accepts only
the Supervisor Ingress source"] N --> DIR["Direct editor, port 80/tcp,
mapped to host 1880,
Supervisor auth by default"] N --> EP["Flow HTTP root /endpoint,
guarded only by http_node's
own Basic Auth, if you set it"] N --> ST["Static content under
http_static, files only"] N --> DASH["Optional FlowFuse Dashboard 2,
routes served by Node-RED itself"]
/endpoint is the one to read twice: it is the only door in this picture with no Supervisor authentication behind it at all, so whatever http_node's own Basic Auth is set to is the entire lock.Think of the shop's front door, the loading dock round the back, and a vending machine bolted to the outside wall. The Supervisor guards the front door and checks every badge. Nobody checks the vending machine unless you personally fit a lock to it — that vending machine is /endpoint, and http_node's Basic Auth is the only lock it will ever have.
What the manifest grants is a ceiling, not a to-do list
hassio_api=true with hassio_role=manager hands the Add-on powerful Supervisor capabilities. homeassistant_api=true reaches the HA Core API. auth_api=true lets the wrapper use Supervisor authentication. host_network=true widens what the Add-on can reach on your local network, and uart=true opens serial hardware access. None of that is a requirement to use all of it — it is the ceiling a misbehaving or malicious node could reach. Never write a Supervisor or HA token into a flow or into Debug output, and keep the list of additional nodes, outbound destinations, and writable directories as short as the job actually needs.
Flows can write to homeassistant_config:rw, media:rw, and share:rw. Allowlist the exact filenames and paths you use, to rule out path traversal, accidental overwrites, and a secret ending up in a shared directory by mistake. UART and Modbus writes can move something in the physical world — validate them in an isolated environment with no hardware attached, every time.
Packages and init_commands run inside the trust boundary
At startup the wrapper installs system_packages, then runs npm install --omit=dev --omit=optional for npm_packages, and finally passes each line of init_commands to eval. A failure at any of those three stages aborts the whole startup. Because init_commands runs at every restart, not just the first one, keep it empty unless every command has gone through change review, is repeatable, and contains no secret.
A few categories of pinned package are worth flagging by what they do, whatever you name them: one hashes HTTP Basic Auth passwords for the wrapper — a hash is not a publishable password. Some can control devices or send outbound messages directly, and belong disabled during any restore drill. Others manage database credentials, query scope, and retention on their own. A few touch OT or serial hardware, where the risk compounds with uart=true. And a couple only read untrusted content or probe the network, where the control that matters is limiting sources, sizes, and destinations — not the package name.
“It's slow” is five different problems wearing the same word
“Node-RED is slow” could mean the editor is drowning in Debug output, a message got too large somewhere, a sequence buffer keeps growing, external I/O is stuck waiting, or V8 is garbage-collecting old space too often. It does not automatically mean you need more memory. The fix is the same one this whole part keeps returning to: find the layer before you touch a setting.
| Layer | Example symptom | Check first | Do not do first |
|---|---|---|---|
| Supervisor / Add-on | A startup loop, or initialization that aborts | Startup time, the first error, recent option changes | Clear settings, or hide the first error under repeated restarts |
| Node.js / runtime | Frequent GC, event-loop latency, or the process exiting | Memory trend, message rate, the old-space setting | Set the heap size close to the host's total RAM |
| Flow / node | Duplicate messages, buffer growth, an unknown node | Node status, Catch output, message shape, Deploy time | Delete an unknown node, or run a Full Deploy to see what happens |
| External I/O | Timeouts, reconnects, delayed data | Destination status, request duration, retry rate, queue depth | Run unauthorized probes, or disable TLS verification to test faster |
| Reverse proxy / auth | Ingress opens, but one endpoint fails | Ingress vs. direct access, port, path, TLS, and auth, checked separately | Turn on leave_front_door_open |
flowchart TD
A["Something feels slow or is failing"] --> B{"Did it start right after
an option change or upgrade?"}
B -->|"yes"| C["Supervisor / Add-on layer:
first error, recent option changes"]
B -->|"no"| D{"Is memory or CPU
climbing over time?"}
D -->|"yes"| E["Node.js / runtime layer:
GC frequency, message rate,
the old-space setting"]
D -->|"no"| F{"Is it one specific
flow or node?"}
F -->|"yes"| G["Flow / node layer:
status, Catch output,
buffer depth"]
F -->|"no"| H["External I/O, or the reverse
proxy and auth layer:
check both separately"]
Debug output and message cloning
A Debug node should show one sanitized field, not the whole message object. The Add-on's settings template sets debugMaxLength to 1000 characters — but that only truncates the display. It proves nothing about whether the full object upstream was ever cloned or retained. Node-RED can clone a message every time it branches, and large nested objects, Buffers, and many branches all make that more expensive. Images, audio, large JSON values, and a live msg.req or msg.res do not belong in routine Debug output or in context, ever.
RBE: a stateful filter, not a stateless conversion
RBE — shown as Filter in the palette — passes a message through only when a value changes, or when a numeric deadband or narrowband rule allows it. It keeps the previous comparison value per topic, which makes it stateful, not a pass-through. You choose the property to compare, typically msg.payload, and the topic property. A deadband comparison expects a number, and because the node reaches for parseFloat under the hood, reject strings and objects before they get anywhere near it rather than trusting permissive parsing to do the right thing.
msg.payload — the usual property RBE compares against its stored previous value.msg.reset — forces RBE to forget its stored value for that topic and treat the next message as new.Test it against an unchanged value, a value right at the threshold, a value past it, a different topic entirely, a msg.reset, and what happens to the comparison state after a restart or a Deploy. That last case matters more than it looks: whether the very next message gets through can depend on state you never explicitly set, so nothing downstream should assume a particular prior state that was never actually tested.
Delay, Trigger, Join — give every wait a bound
msg.parts correct, and cap the number of waiting sequences, items per group, and wait time. nodeMaxMessageBufferLength is the Add-on's own limit for this, and its default of 0 means unlimited — not evidence that a limit exists.Context and external I/O
Node, flow, and global context can live in memory or in a named localfilesystem store. Neither is the place for unbounded arrays, whole event objects, or binary payloads, and a disk-backed store is not guaranteed to write durably on every single assignment — decide the limits on keys, size, retention, and write frequency before you rely on it. HTTP, MQTT, WebSocket, TCP, UDP, serial, files, databases, and HA's Get History or Poll State are all external I/O in this sense: give each one a timeout, a result-size limit, a retry limit, and a backoff policy. Limit the time range and result count on Get History specifically, and reach for an event subscription instead of Poll State wherever the two would do the same job.
Safe Mode is neutral gear, not the engine off
The Add-on option safe_mode: true adds the --safe flag at startup. Node-RED still starts — it just suppresses the initial startup of your flows, giving you a window to fix something before it runs again. What it does not do is stop a Deploy from starting the flows in that deployment, even while safe_mode is still set to true. It is not Home Assistant's own Safe Mode, it does not stop the Add-on, and it changes nothing about editor or admin authentication — Ingress and direct-access rules still apply exactly as before.
Safe Mode is putting the car in neutral, not switching the engine off. The engine keeps running. Nothing moves until you press a pedal — and Deploy is the pedal. Press it, and whatever is in that deployment goes, safe_mode or not.
sequenceDiagram participant Op as Operator participant AO as Add-on, safe_mode true participant NR as Node-RED runtime Op->>AO: restart with safe_mode true AO->>NR: start with --safe NR->>NR: suppress the initial startup of flows Op->>NR: click Deploy NR->>NR: start every flow in that deployment
safe_mode actually buys you: flows stay stopped through the restart. The last two arrows happen regardless — a Deploy from the operator still tells the runtime to start what is in it, which is why disabling side-effect nodes has to happen before that click, not after.Before the first Deploy after any restore or repair, disable every node with a network call, an HA action, a write, or a schedule attached — or disconnect the wire leading into it. Do not Deploy first and clean up second; a Deploy that already ran cannot be undone by a check that comes after it.
Deploy scope is not a performance switch
Node-RED 5.0.2's editor offers three Deploy scopes: Full, Modified Flows, and Modified Nodes. Which one you pick decides which flows or nodes actually restart, and that in turn affects timers, open connections, context lifecycles, and any message mid-flight. Understand what depends on the thing you are changing before you choose the narrowest scope — if a shared config node's blast radius is not clear to you, do not reach for Modified Nodes just because it looks faster.
max_old_space_size limits one drawer, not the whole cabinet
This Add-on option is an integer in MB. When it is set, the startup script exports exactly one line:
NODE_OPTIONS=--max_old_space_size=PLACEHOLDER_MBmax_old_space_size is turning up the size of one drawer in a filing cabinet, not renting a bigger office. Old space is one section of the V8 heap. The young generation, native add-ons, Buffers, and general overhead all live in other drawers this setting never touches — and the cabinet still has to fit in the room, which is the Home Assistant host's actual RAM.
- Record process and container memory, message rate, latency, restarts, and queue depth first — before changing anything.
- Fix unbounded Debug output, oversized messages, unbounded context, and unbounded Delay/Trigger/Join buffers before you even consider this setting. More old space does not fix a leak.
- If a real need remains, pick a conservative value, change only that, and repeat the same load in an isolated environment.
- Watch GC behavior, latency, and host headroom. No improvement means roll back — not raise the number again.
Log level, and what not to paste into a ticket
log_level accepts trace, debug, info, notice, warning, error, and fatal, and the option can be left unset. The official docs recommend info for normal use; the wrapper maps warning onto Node-RED's own warn. Turning verbosity up can print payloads, URLs, headers, entity or device IDs, MQTT topics, file paths, and stack traces straight into the log — use a higher level only for a defined window, then put it back. When you do share a log, keep versions, timestamps, node types, error categories, and a correlation placeholder, and strip everything else: tokens, passwords, Authorization and Cookie headers, encrypted credentials, private keys, webhook IDs, location data, real hostnames and IPs, full payloads, and anything visible in a screenshot's edges.
Six startup stages, and each one has its own way to fail
Before you upgrade anything: finish an Add-on backup and an isolated restore drill, with credential_secret or the Project secret stored separately from both. Export a sanitized flow inventory — Add-on, Node-RED, and HA WebSocket versions, every extra npm and system package, and the optional, not-bundled FlowFuse Dashboard 2 if you use it. Its own metadata calls for Node 14 or newer and Node-RED 3.0.0 or newer, and that compatibility line is not a substitute for testing it yourself in your own environment. Read the target release's notes and breaking changes. Disable HA, network, write, and scheduling nodes, and decide your stop condition and rollback version before you start — not while you are already mid-upgrade.
| Stage | Behavior in 22.0.1 | Safe response to a failure |
|---|---|---|
| Data initialization | A fresh /config may migrate data from /homeassistant/node-red; otherwise settings, flows, and node directories are created new | Inventory both directories before touching anything. Never move files by hand — find the first migration error instead |
| Theme migration | An old dark theme is renamed to dark-modern | That is a name change, not flow corruption — do not misdiagnose it as one |
| Conflicting packages | The wrapper tries to remove node-red-contrib-home-assistant, node-red-contrib-home-assistant-llat, and node-red-contrib-home-assistant-ws | This is scripted wrapper behavior — do not run your own removal command. If it fails, keep the log and the package inventory |
| Custom packages | Alpine system_packages install first, then npm npm_packages; any failure aborts startup | Find the first package, repository, or architecture error, and roll back only the most recent option change |
| Custom commands | Each line of init_commands is passed to eval at every startup; any failure aborts startup | Disable the most recently added command and return to known-good settings. Never put a secret in a command or a log |
| Runtime flags | safe_mode becomes --safe; max_old_space_size becomes NODE_OPTIONS | Confirm each option exists with the right type, and do not confuse the manifest's init=false with runtime Safe Mode |
After an upgrade, stay in Safe Mode and verify the editor, node types, credential decryption, settings, and package loading — all without deploying. Only once every side-effect node is disabled or disconnected does a Deploy of the smallest pure-data flow become the next, separately approved, step. If rollback becomes necessary, it means more than swapping the image back: stop making further changes, preserve sanitized logs from the failed version, and use the Home Assistant Add-on's own restore procedure to land on a matching backup, version, and secret together. Do not feed data a newer version already migrated back into an older runtime, and never hand-edit a node's type or version field inside flow JSON.
The failures that actually happen, matched to what to check first
| Symptom | Likely cause | What to do |
|---|---|---|
| Credentials don't work after a restore | Wrong secret, wrong version, or a Project boundary crossed | Keep flows disabled. Verify the backup's version and Project boundary against the secret you have. Don't keep changing credential_secret hoping one attempt works — if no matching secret exists, reissue the external credentials instead |
| Ingress opens, but /endpoint/ has no authentication you expected | You're thinking of Supervisor auth, which /endpoint never had | Confirm whether you're on Ingress or the direct port, then test with a read-only, no-real-data request. Check http_node, TLS, and the port mapping — never reach for leave_front_door_open |
| Unknown nodes appear after a restore | A package version wasn't rebuilt, or doesn't match the inventory | Don't deploy or delete it. Match it against the pinned-package inventory, rebuild the excluded node_modules in isolation, and verify source and version before it loads |
| The Project remote or its credentials fail | A mismatched secret, remote identity, or access scope | Preserve the local data and Git status as-is. Verify the Project secret and remote identity. Don't force-push, reset the secret, or reinitialize the Project |
| You suspect a token or key leaked | Exposure through a log, a backup, Git, or a shared directory | Isolate the Add-on first and preserve a de-identified timeline, then rotate the credential at its source. Deleting the Debug message is not the fix — check backups, Git, and shared directories for the full exposure |
| Add-on startup stops at a custom package or command | A system_packages, npm_packages, or init_commands failure | Find the first error in the log, compare it against recent option changes and the architecture, and roll back only the latest change — not an ad hoc shell command |
| An HA node reports disconnected or Unauthorized WebSocket | The connection mode setting, or a version below 0.80.3's own floor | Confirm “I use the Home Assistant Add-on” is set correctly in the HA Server config. Check HA 2024.3+, Node-RED 3.1.1+, and Node 18.2.0+ — don't copy tokens or send a test Action to debug it |
| An HTTP node or Dashboard route 404s | Confusing Ingress with the direct, separately mapped port | HTTP nodes need their own mapped port and live under /endpoint/. Check the port, the TLS pair, and http_node authentication — not TLS verification or leave_front_door_open |
| A TLS handshake or certificate error | An expired, mismatched, or wrongly placed file | Check filename, validity period, subject, and chain under /ssl only. Never expose a private key or disable verification to move faster |
| Join or Wait Until emits nothing, and memory keeps growing | An unbounded sequence, timeout, or pending-key set | Sample msg.parts, group size, timeouts, and pending keys. Stop new input and preserve evidence — don't inject more test traffic or clear production context directly |
| Repeated restarts hid the real first error | Only later timeouts survived, not the original failure | Stop restarting. Treat the earliest first error and the last known good time as primary evidence, and label anything after that as secondary. Don't guess the root cause from what's left |
| Unknown nodes or schema warnings appear after import or upgrade | A pinned package version wasn't matched, or a legacy node type has been deprecated | Stay in Safe Mode. Don't Deploy, delete nodes, or edit their type manually. Identify the pinned package from the backup inventory and restore a compatible version in an isolated environment first. The legacy HA Entity node is deprecated — handle it only through the migration procedure for the pinned version, and don't create a new one |
| An upgrade report or escalation is required | The issue survives isolated troubleshooting and needs a maintainer | Stop making repeated changes. Give the Add-on or node-package maintainer pinned versions, the architecture, a reproduction timeline, the first error, whether the problem remains with side effects disabled, a minimal sanitized flow, and the one change already attempted. Never attach tokens, credentials, private keys, real endpoints or IDs, full payloads, browser storage, or unredacted screenshots |
The ones that come up again and again
Can I rotate credential_secret on a schedule, like an access token?
Can a private Git repository for Projects replace an Add-on backup?
Does logging into Ingress protect HTTP, Dashboard, and MQTT too?
/endpoint/, static content, Dashboard routes, WebSockets, and MQTT are separate boundaries, each with its own authentication story. Set each one up on its own terms rather than assuming one login covers the rest.How do I run a restore drill without triggering real smart-home actions?
Can I set max_old_space_size to my host's total RAM?
Does Safe Mode stop the whole Add-on?
--safe and only suppresses the initial startup of flows. Any Deploy — even while safe_mode stays true — can still start the flows in that deployment. Disable or disconnect every side-effect path before that first Deploy, not after.Does truncating Debug output reduce its actual memory cost?
debugMaxLength limits the display only — it says nothing about whether the full object was cloned or retained upstream. Reduce the message at its source, and separately limit both how often Debug fires and which fields it shows.Can I just delete and recreate an unknown node after a restore?
My HA node says disconnected right after an upgrade. Where do I look first?
Why does the restore drill insist on an isolated environment at all?
During an incident, the secrets-manager reference does not match the backup time. Which key should I try first?
Where to go from here
You now have a Node-RED install that talks to Home Assistant, and the habits to keep it that way
This closes the series: safe foundations, Home Assistant in practice, and the engineering and operations that keep both running when something goes wrong instead of when everything is fine. If an earlier part sent you back here — a failed restore, a slow flow, an upgrade that would not settle — the fix is almost always in one of the tables above. Keep this part bookmarked longer than any other; it is the one you come back to.
Open the full guidePart 12 of the WoowTech Complete Node-RED Guide series on the Apporo blog.
Adapted from the WoowTech Complete Node-RED Guide, produced by WoowTech and released under CC BY 4.0. This adaptation is published by Apporo under the same licence.
Light · Air · Water · Control · apporo