Every intent starts as a proposal, and the Action node stays off until you turn it on yourself
Up to here the series has been about reading Home Assistant — listening for events, querying state, calling one Action you already trusted. This part is about deciding when not to call it. A hallway light, a door left open, a room that is getting warm: three ordinary situations, and three flows that only ever produce a candidate. Every Home Assistant listener, query, wait, and Action node in this part stays disabled, wired to nothing but a redacted Debug, until you have proved — with synthetic data, then with a read-only test entity — that the flow says no as often as it should.
A flow that turns a light on is easy. A flow that knows when not to is the actual job
A motion sensor can change state several times inside a few seconds, or it can sit on clear for a while and then flicker back to active the moment someone reaches the bottom of the stairs. A wall switch, Home Assistant, another automation, or a device recovering from a dropped connection can all change a light's state without your flow ever being asked. Feed a single fixed Delay into that situation and you get a race: an old turn-off message arrives after new motion, and the room goes dark on someone. Give every motion event its own wait instead, and you can end up with several pending timers arguing over the same light.
The scenarios below are deliberately non-critical, and deliberately disconnected from anything real while you are building them.
PLACEHOLDER_LIGHT_ENTITY and PLACEHOLDER_MOTION_ENTITY — placeholders, not identifiers for anywhere real.The same discipline runs through the notification and climate side of this part: PLACEHOLDER_NOTIFY_ACTION, PLACEHOLDER_NOTIFY_TARGET, PLACEHOLDER_CLIMATE_ACTION, PLACEHOLDER_CLIMATE_ENTITY, and PLACEHOLDER_OPENING_ENTITY must not be swapped for real values and deployed directly. They exist so the design can be reviewed field by field before a single real device is anywhere near it.
Event, state, intent, and action are four different questions
It is tempting to build straight from “motion detected” to “light on,” with one Switch node in between. That collapses four separate questions into one, and each of the four has its own answer and its own failure mode.
| Layer | Example data | What it must not assume |
|---|---|---|
| Event | A state_changed event arrives for the sensor entity | That an event necessarily means someone is still present |
| State | Active, clear, unknown, unavailable | That every binary sensor displays the same text for the same condition |
| Intent | Propose light-on, propose light-off, cancel light-off | That the device has actually carried out the intent |
| Action | Call an approved action to turn the light on or off | That the call necessarily succeeds, or that every service returns the same response shape |
Three mechanisms sit inside that Intent layer, and conflating them is the single most common way a “working” flow starts misbehaving under real conditions. Debouncing requires an input to stay stable for a period before you accept the transition at all. A cooldown suppresses another intent of the same kind for a period after one has already been generated. A delayed light-off is a business-rule wait — give the room a few minutes of grace after it reads clear, in case that reading is wrong or someone is about to walk back in. Calling all three “a Delay node” makes it impossible to tell, six months from now, whether a new event was supposed to reset the wait, queue behind it, or run alongside it.
Debounce is waiting for a door to stop swinging before you trust whether it is open or shut. A cooldown is refusing to shout the same warning twice in one breath, even though the door swung again. A delayed light-off is a different decision entirely: once the room reads clear, hold off for ten minutes before switching anything off, in case that reading turns out to be wrong. Treat all three as one Delay node and you get a light that goes dark on a guest who is still standing in the hallway, because the flow only ever built the third behavior and quietly expected it to cover the other two as well.
Events: state for one transition. Trigger: state when you need conditions
Before wiring anything, write down — in an offline document, not in the flow — the values your sensor integration can actually produce, then map each one to exactly three buckets: ACTIVE, CLEAR, and INVALID. Treat unknown, unavailable, null, and a missing new_state as INVALID, never as CLEAR. An invalid reading is not a clear room; it is a room you know nothing about.
Events: state is the narrower node, and it is easier to review when the requirement is simply “for this one entity, accept a transition from a known value to active or clear, optionally after it holds for a set duration.” Trigger: state earns its place when you need allowed and blocked outputs, several custom outputs, or the ability to feed it test messages. In HA WebSocket 0.80.3 it can filter entities by exact match, list, substring, or regex, and it supports Output on connect, Enable input, and synthetic test input on top of that. More flexibility asks for tighter discipline:
- Prefer exact matching in production. Substring or regex matching can silently widen scope the day someone adds a new entity that happens to match the pattern.
- Keep Output on connect disabled by default, so a Node-RED restart is never read as fresh motion.
- Leave Enable input off by default. If you turn it on for offline testing, accept only synthetic events with a fixed schema, then switch it back off.
- Route a failed condition to a diagnostic, blocked branch — never read blocked as “turn the light off.” It only means the current intent was not allowed through.
- Let State Type conversion touch only known values. Isolate
unknownandunavailableas raw strings first; do not let the node infer a number or a boolean from either.
Events: state is a single, fixed question: did this one entity change the way I said it would? Trigger: state is a form with extra boxes ticked — conditions, several outputs, a way to feed it test data. A form with more boxes is not automatically the better choice. Every extra box is one more setting to get right, and Output on connect left switched on is the box almost everyone forgets, because it quietly reports “connected” as if it meant “motion detected,” the instant Node-RED comes back up.
entity_id, old_state, and new_state to simulate an event, but that only validates the node's own conditions and outputs. It proves nothing about how the Home Assistant integration, the wireless transport, or the physical light will actually behave. Use placeholder IDs for every synthetic event, and keep the node disabled until isolated testing has been approved.flowchart TD A["Need conditions, several outputs,
or test-event input?"] -->|"no"| B["Events: state"] A -->|"yes"| C["Trigger: state,
exact-match filtering"] B --> D{"Normalized value?"} C --> D D -->|"ACTIVE"| E["Cancel pending light-off,
check cooldown and override,
propose light-on"] D -->|"CLEAR"| F["Debounce, then start
one pending light-off"] D -->|"INVALID"| G["Cancel pending light-off,
redacted diagnostic only"]
Timestamps from either node can still carry source latency — a reading that left the sensor two minutes ago is not the same as one that just happened, even if both arrive in the same second. The normalization layer should keep both the event time and the receipt time, and reject anything that moves noticeably backward or exceeds a freshness threshold you have set. Keep Debug output to an anonymized source, the normalized state, and a latency band; do not let it retain the complete old_state and new_state objects.
A generation counter, a Do Not Disturb sign, and a recheck that trusts neither
In Node-RED 5.0.2, a fixed Delay creates an independent timeout for every ordinary message, a new message does not automatically reset an earlier one, and there is no per-topic reset built in. msg.reset clears every pending message in that Delay node at once — all of them, not just the one you meant. For debouncing, use the core Trigger node's extend-delay option with bytopic instead. If a fixed Delay is what implements the wait after CLEAR, send msg.reset explicitly before replacing the candidate, then send the one candidate you actually want, and give each area its own Delay so one area's reset cannot silently clear another's. A deploy or a process restart can still erase an in-memory wait, so “no timer is currently running” must never be read as permission to turn the light off.
Wait Until is the alternative when what you are actually waiting for is a state condition rather than a fixed duration. After it receives input, it listens to the specified entity until its property matches your comparator and value, with a timeout and separate success and timeout outputs. With Check against current state enabled, an immediate check is available for exactly one entity. Each Wait Until node holds only one active slot — a later input cancels the existing timer and replaces the saved message and config rather than queuing behind it. msg.payload can also override the entity, property, comparator, value, and timeout, which is exactly why a safety-conscious flow enables Block Input Overrides: it keeps an upstream message from quietly retargeting what the wait is even watching. A timeout of 0 creates no timer at all; it is not a finite wait.
Give every candidate light-off a monotonically increasing generation, or request ID. A new ACTIVE advances the generation. When a wait ends, discard its message immediately if the generation it carries is no longer current. That closes the race where a cancellation and an expiry arrive at almost the same moment — but a generation number is not a device state, so the final Current State recheck still has to run regardless of which generation won.
sequenceDiagram participant S as Motion sensor participant F as Flow participant W as Pending wait S->>F: CLEAR, debounce done F->>W: start wait, generation 1 S->>F: ACTIVE arrives mid-wait F->>W: msg.reset, generation 2 W->>F: generation 2 expires F->>F: recheck current state before acting
A manual override is a Do Not Disturb sign, not a power switch. It carries a source, the time it went up, and the time it comes down on its own. While it is hanging there the flow keeps watching — it just is not allowed to act on what it sees. And nobody assumes it is fine to walk straight in the instant the sign comes down, either: the flow re-evaluates from scratch instead of rushing to reverse whatever the person changed while the sign was up.
An override value should carry its source, its creation time, and its expiration time — never a person's name or location. Turning the light on by hand can hold off automatic light-off for the override's duration; turning it off by hand should keep an activity signal from switching it straight back on. Expiration itself triggers nothing — it only resumes evaluation, and the flow still waits for the next trusted event or queries state again rather than acting the instant the clock runs out. After a Node-RED restart, if override data is missing or has expired, the conservative default is: do not turn the light off automatically. The same shape of rule governs the climate side of this part later on — a manual thermostat change is a first-class flow state there too, not an error to route around.
Persistent context carries its own obligations: old data versions, an expiration you have to honor, and backup privacy. In-memory context loses every candidate and override the moment Node-RED restarts. Neither choice is automatically safe — the restart branch belongs in your state machine explicitly, drawn out rather than assumed away.
Offline review, synthetic data, one read-only node, then a walkthrough that still calls nothing
Both the motion-lighting design and the notification and climate design move through the same shape of test plan, and neither one lets a real Action fire before the last stage — and even the last stage keeps it disabled.
flowchart LR A["Stage 0
Offline review"] --> B["Stage 1
Synthetic data,
core nodes only"] B --> C["Stage 2
Read-only observation,
one HA node enabled"] C --> D["Stage 3
End-to-end walkthrough,
Action still disabled"]
-
Stage 0
Offline review
Inspect the flow JSON, the node settings, and the wiring without opening a Node-RED connection at all. Build a test matrix covering active, repeated active, clear chatter, clear followed by active, unknown, unavailable, a missing
old_state, a restart, and a manual override. Every ID in the flow must be a placeholder at this point. -
Stage 1
Synthetic data, core nodes only
Use a manual Inject to send normalized states — not the automatic repeat or once options. Debug shows only
scenario,decision, andgeneration. Confirm that every case in your matrix produces at most one intent. There is no Home Assistant Server config anywhere in this stage, and no external side effect is possible. -
Stage 2
Read-only observation
With separate approval, enable exactly one Events: state or Trigger: state node against an isolated Home Assistant test entity. Current State and Wait Until can stay disabled; the Action must stay disabled. Compare real event ordering, real chatter intervals, and real unavailable behavior — without recording detailed occupant activity anywhere.
-
Stage 3
End-to-end walkthrough, Action still disabled
Configure the Action node with
PLACEHOLDER_HA_ACTIONandPLACEHOLDER_ENTITY_ID, keepd: true, and enable Block Input Overrides. Wire the upstream path only into a “prepare action” Debug — never into the Action itself. Do not treat the example's message shape as a universal schema for every lighting service.
The notification and climate design in the rest of this part runs the same shape of plan, just with an eighth step folded in: matrix tests that specifically cover opening jitter, recovery, duplicate events, notification bursts, non-numeric values, unit errors, and simulated Action failures, using nothing but manual Inject and Debug nodes with no external connection at all.
One reminder per opening, not a running commentary
A door or window sensor turning on might mean ventilation, a door left open, sensor jitter, or a stale state replayed after a reconnect. Send a notification for every one of those and you train people to swipe it away without reading it. Build the candidate object in the pure message layer first, and do not hand it to an Action directly:
{ "kind": "OPENING_STILL_OPEN", "anonymous_area": "PLACEHOLDER_AREA_ALIAS", "event_generation": "PLACEHOLDER_EVENT_GENERATION", "severity": "informational", "observed_at": "PLACEHOLDER_ISO_TIMESTAMP" }A transformation layer then builds the real payload according to the official schema for whichever notification action you have actually approved — field names and advanced data are not portable between integrations, so do not assume they are. Message content should never reveal that “no one is home,” a precise address, an opening's real ID, or a person's name.
- Fix the action to
PLACEHOLDER_NOTIFY_ACTION, with no dynamic override frommsg, and Block Input Overrides enabled. - Fix the target to
PLACEHOLDER_NOTIFY_TARGET— never a real mobile-app device ID, person, area, label, or group. - Build the data from an allowlist. Do not merge the whole
msg.payload, context, or original event straight into it. - A Queue policy is not a rate limit. Items queued during a disconnection can all land at once after recovery; notification flows should generally discard an expired intent instead of choosing Queue all.
- Action output only tells you the node's own call path finished. It does not prove every notification service returns the same content, or that an endpoint actually displayed anything.
A Queue setting on the Action node is about delivery, not restraint. Queue all is a promise to eventually send everything that piled up while the connection was down — exactly the wrong promise for a reminder that “the window is still open,” three hours after someone already closed it. What actually stops a notification storm is deduplication and a rate limit sitting upstream of the Action, deciding what even reaches the queue in the first place.
| Rule | Example | When it is exceeded |
|---|---|---|
| Event deduplication | At most one reminder per opening generation | Discard the duplicate, increment an anonymous counter |
| Event cooldown | No repeat reminders while the opening stays open | A new generation starts only after a confirmed recovery |
| Area rate | Cap notifications per anonymous area, per time window | Do not send; mark rate_limited |
| Global rate | Protects recipients when several sensors fail at once | Trip the notification Action's circuit breaker; keep a summary only |
| Expiration window | An old event exceeds the tolerated age after a reconnect | Do not resend; record a recovery summary only |
Never retry a failed notification indefinitely. Set a maximum attempt count, a backoff, and a total time limit, and keep retries subject to the same deduplication key and the same global limit. Safety-critical alerting belongs on a dedicated, monitored channel — the general notifications described here carry no life-safety guarantee of any kind.
The opening's own life cycle
Doors and windows are usually binary sensor entities, but confirm the actual state and device class for your own integration rather than assuming on means open because the icon looks that way. Observe one known-open and one known-closed reading in an isolated environment before you write the mapping down.
flowchart TD A["Opening sensor reports
a new state"] --> B{"Normalized value?"} B -->|"CLOSED to OPEN"| C["Debounce, then
start a generation"] C --> D{"Still OPEN when
the wait ends?"} D -->|"yes"| E["Check dedup and
rate limit"] E --> F["Candidate reminder,
disabled Action"] D -->|"no, closed early"| G["No reminder sent"] B -->|"OPEN to CLOSED"| H["Cancel any
unexpired candidate"] H --> I["Recovery notice, only if
a reminder went out earlier"] B -->|"unknown or unavailable"| J["Cancel control intent,
route to exception diagnostics"]
If it closes within the required duration, no “still open” reminder goes out. If it closes after a reminder already fired, your product rules decide whether a recovery notice is warranted, and that notice should share the original event's correlation key so it is not read as a fresh alert. After a Node-RED or Home Assistant restart, query the current state and the freshness of last_changed first — do not replay an old generation out of context.
Opening events can reveal arrivals, departures, and daily routines on their own. Keep Debug output to the anonymous area, OPEN/CLOSED/INVALID, a duration band, and the decision — never an opening's real name, a user ID, a precise timeline, or raw attributes. Calendar, Zone, Tag, and Sentence can each add a useful, conservative condition, but every one of them widens what the flow can see about a household:
- Calendar can produce an “approved time window” condition, but its roughly 15-minute polling interval and vulnerability to last-minute changes make it unfit as the sole security condition.
- Zone can produce an anonymous enter/leave condition, but location drift or stale coordinates must never directly dismiss an opening reminder on their own.
- Tag can provide a manual-confirmation input, but a tag scan is not identity verification and must not directly trigger an unlock or a climate change.
- Sentence can accept a “remind me later” intent, but its response mode sends an external reply, so keep it disabled and never assume the speech recognition or the reply will succeed.
Six checks a reading has to pass before it is even allowed near a threshold
A temperature reading above a threshold might reflect a genuine warm room, or it might reflect a string-parsing error, mismatched units, stale data, or a sensor that has simply gone unavailable. Controlling equipment straight off one reading can waste energy, cause rapid cycling, or fight a setting someone just chose on purpose. Six gates run before the value is even compared to anything.
flowchart LR A["Value exists,
not unknown/unavailable"] --> B["Converts to a
finite number"] B --> C["Unit matches
the flow's config"] C --> D["Inside a reasonable
range for the sensor"] D --> E["Fresh enough,
not a stale reload"] E --> F["Source is on the
exact-match allowlist"] F --> G["Only then:
compare to the threshold"]
Do not flip control on a single threshold. A high reading can generate a “cooling needed” candidate, which then clears only once the value drops below a separate, lower recovery threshold and stays there for a defined duration. That hysteresis gap is what keeps the equipment from clicking on and off every time the reading drifts across one line. The exact values depend on the space, the equipment, health requirements, and your own energy policy — there is no number here meant to be copied into a real thermostat.
Passing the six gates only earns the reading a comparison. A candidate control intent still needs every one of these before it reaches the disabled Action:
- Every relevant opening reads a known CLOSED state; any OPEN or INVALID reading blocks the climate candidate outright.
- The climate entity exists and is not unavailable, and its current mode and attributes match what the integration documents.
- No manual override is active — if someone just changed a setting, the flow does not counteract it.
- A similar intent is not still in cooldown, and the equipment is not at risk of short-cycling.
- The Action and its data have been checked against what the entity actually supports; assume nothing about a generic service response.
NO_ACTION_INVALID_DEPENDENCY — nothing else.A manual override here carries the same shape of contract as the one in motion lighting: an anonymous control scope, a direction, a creation time, an expiration time, a source category, and a hard maximum lifetime. While it is active, the flow may quietly record “the candidate that would have fired without the override,” purely to help tune the rules later — but it must not send an Action. Query again once the override expires; do not replay whatever candidates piled up while it was active. If Node-RED cannot confirm override state after a restart, the fail-closed default is the same one as always: no control.
Real Action calls use PLACEHOLDER_CLIMATE_ACTION and PLACEHOLDER_CLIMATE_ENTITY, and the data contains no real setpoint — PLACEHOLDER_APPROVED_VALUE only marks a value that is still awaiting approval. Any real equipment operation needs its own separate validation, in a non-production environment, against one limited target, with a manual abort available the whole time.
None of the eleven nodes below is a “safer variable.” Several can accept operations from Home Assistant, or push state back into it, so keep all of them disabled during testing.
| Node | Role here | Safety boundary |
|---|---|---|
| Events: calendar | Candidate for an approved time window | Polling delay and schedule privacy; never the sole security condition |
| Sentence | Manually postpone a reminder | The utterance and device/response ID are sensitive; a response has a real side effect |
| Tag | Manual confirmation input | A tag scan is not strong authentication |
| Zone | Anonymous location condition | Coordinates are sensitive and drift; never controls a lock or climate directly |
| Select entity | Override interface, limited options | Accept only listed options; keep the node disabled |
| Binary Sensor entity | Exposes an anonymous exception flag | Never coerce an unknown reading into false |
| Button entity | Manual candidate-confirmation press | A press is not authorization; never wire it straight to an Action |
| Number entity | Exposes a constrained candidate threshold | Validate the number, its bounds, and its unit before use |
| Sensor entity | Publishes an anonymous health summary | Never publish sensitive raw data through it |
| Switch entity | Time-limited intent to enable automation | Defaults to enabled with no persisted value — add a fail-closed startup gate |
| Text entity | Constrained, non-sensitive annotation | Validate length and content; never carry secrets or arbitrary actions |
Keep the exception envelope itself just as plain, whichever of the two flows raised it:
{ "category": "INVALID_STATE_OR_VALUE", "source_alias": "PLACEHOLDER_SOURCE_ALIAS", "observed_at": "PLACEHOLDER_ISO_TIMESTAMP", "decision": "NO_ACTION", "detail_code": "PLACEHOLDER_NON_SENSITIVE_CODE" }No raw state, attribute, location, schedule, utterance, notification target, or device ID belongs in that envelope. Catch and Status nodes can gather node errors and status safely, but they must never build a “resend the Action on error” loop; once a counter crosses its threshold, trip the circuit for that path and bring it back only after a manual review.
- Notification integration unavailable: do not queue indefinitely, and do not fall back to an unapproved target.
- Climate equipment unavailable: do not control it, and do not pretend it reached the setpoint.
- Opening sensor unavailable: never treat that as closed; block the climate candidate.
- Entity config or the custom integration not loaded: disable the exposed-entity path without touching the core no-action policy.
- Reconnect or restart: clear expired candidates and query again; never send an old notification or replay an old control.
The symptom you see, and the check that actually explains it
| Symptom | Likely cause | What to do |
|---|---|---|
| One activity produces several light-on intents | Events: state and Trigger: state are both wired in, or property-only updates are being observed too | Disable every Action, count by event time and anonymized source, then add deduplication and a cooldown |
| A light-off intent still lands while someone is moving | A new ACTIVE did not actually send msg.reset, or two areas share one Delay, or the generation is stale |
Check reset is sent on every ACTIVE, give each area its own Delay, and recheck state before the wait expires |
| Wait Until waits forever | The entities filter, property path, raw state string, or comparator does not match what the entity actually sends | Set a finite timeout and wire a separate timeout branch. A timeout means the condition was not observed — it does not mean light-off |
unknown or unavailable is treated as clear |
Switch-rule order or type conversion let it fall through to a default | Intercept INVALID first and cancel the candidate; it must never reach a default CLEAR branch |
| The light turns back on right after a manual light-off | The manual override was not checked before the entry path, or its source was never recorded | Create a time-limited override during which the flow observes activity but never generates an opposing intent |
| Repeated notifications for the same opening event | The deduplication key contains an unstable timestamp, or a persistent OPEN state keeps minting a new key | Keep the generation unchanged until CLOSED, then verify with a synthetic event that one OPEN produces at most one candidate |
| A burst of old notifications after a reconnect | A Queue or retry path retained expired intents through the outage | Add a maximum event age, a global rate limit, and a recovery circuit breaker; record an anonymous summary instead of resending |
| Temperature comparisons look reversed | String-to-number conversion, Celsius versus Fahrenheit, or the hysteresis direction is wrong | Check the conversion, the threshold, and the recovery direction; keep the raw value only for short, controlled diagnostics |
| An open door still produces a climate candidate | The integration's on/off semantics or device-class mapping does not match what the Switch rules assume | Block both OPEN and INVALID explicitly — not just the literal string on |
| The Action node reports no error, but nothing arrives | Node success is being read as delivery success | Keep the node disabled and check the specific notification integration's documentation for its action, target, data, and endpoint state directly |
The ones that come up again and again
Should I reach for a Delay, Trigger, or Wait Until?
bytopic for input debouncing. A separate fixed Delay per area works for the business-rule wait after CLEAR, provided you explicitly send msg.reset before replacing a candidate — ordinary new messages do not reset the timer on their own, and reset clears every pending message in that node. Reach for Wait Until when you need to wait on an entity property and distinguish success from timeout. All three still need a restart strategy and a race-condition strategy on top.Can Trigger: state's blocked output turn the light off directly?
Why does the flow need a Current State query right before light-off?
When is it safe to enable the real Action nodes?
Can an unavailable door or window be treated as closed, just to stop the reminders?
Is it fine to merge notification Action data straight from msg.payload?
Does Queue all protect against missed notifications during a disconnect?
Can a climate Action fire as soon as the temperature crosses a threshold?
Can a Switch entity work as a safe master switch for the automation?
When the sensor becomes unavailable, can I start the light-off countdown?
Can I use a regex for every sensor in an area?
Does Action output mean that a phone displayed the notification or the equipment finished its operation?
Where to go from here
Every intent in this part started as a raw value someone or something reported. Part 8 is about trusting that value at all.
You now have a working shape for a flow that proposes instead of acting: layers kept apart, a debounce distinct from a cooldown, a generation number that survives a race, and four testing stages that never let a real Action fire before the last one. Part 8 turns to JSONata and templates — the expressions that reshape a message before it reaches any of this — and then the Function node, where the same discipline has to survive plain JavaScript.
Open the full guidePart 7 of the WoowTech Complete Node-RED Guide series on the Apporo blog.
Adapted from the WoowTech Complete Node-RED Guide, produced by WoowTech and released under CC BY 4.0. This adaptation is published by Apporo under the same licence.
Light · Air · Water · Control · apporo