Back up the identity, verify it works, then fix what breaks without resetting anything
Twelve parts back you built bridges, filtered and mapped entities, and commissioned them into Apple Home, Google Home and Amazon Alexa. None of that survives on its own. This last part covers the two things that decide whether all of it is still there next month: taking a backup that actually restores, and working through a failure — No Response, a stuck commissioning window, a host that will not stay up — without reaching for a reset before you have to. That order matters. A reset is at the very end of this part for a reason.
Downloading a file and being able to restore from it are two different claims
What can actually be recovered from a Matter bridge is: the bridge config, the entity mappings, the Matter identity and fabric credentials, the plugin enabled / disabled state of each bridge, and a handful of administrative assets. That list matters because of what is missing from it once the identity goes. If the identity is lost you can rebuild the bridge settings from a config backup, but every controller that already commissioned this bridge still has to commission it again — the identity is what a controller recognizes, not the settings behind it. And if you ever start that same identity running on two hosts at once, you get duplicate services and mDNS and session behavior that nobody can predict for you in advance.
Settings → Backup & Restore is where all of this lives: downloadable config backups and full backups, a restore preview, per-bridge selection, an overwrite switch, include mappings, restore identity, internal snapshots, auto backup and retention. A restore can overwrite your current settings and ask the application to restart, so the rule before every restore is the same one you will see repeated through this part: take one more backup of the current state first, so you have an explicit way back.
A config backup is a copy of the rental agreement. It tells you the address, the terms, who lives where. A full backup includes the agreement and a working copy of the front door key. Lose the key and the agreement is still perfectly true — you just cannot get back in without going through the whole business of getting a new key cut and re-registering it with everyone who checks it, which is exactly what recommissioning is.
| What you are actually facing | Evidence you need first | Safe fallback point |
|---|---|---|
| Just a mis-edited setting | A bridge export or config backup, and a diff against it | A selective restore that leaves the identity alone |
| The host or the storage volume failed | A full backup, its version, and a checksum you trust | Stop the old instance, then restore in full |
| A controller shows No Response | Health, sessions, mDNS, and the logs | Fix the network or restart that one bridge — do not reset first |
| The identity is confirmed unusable | Proof the backup will not restore, plus a controller cleanup plan | A factory or disaster reset, as the last resort |
flowchart TD A["What are you actually facing?"] --> Q1["Just a mis-edited setting"] A --> Q2["The host or the volume failed"] A --> Q3["A controller shows No Response"] A --> Q4["The identity is confirmed unusable"] Q1 --> F1["Selective restore,
identity left alone"] Q2 --> F2["Stop the old instance,
then restore in full"] Q3 --> F3["Fix the network or restart
that one bridge, no reset"] Q4 --> F4["Factory or disaster reset,
the last resort"]
Bridge export, config backup, full backup, snapshot — read the label before you trust it
A bridge export JSON, covered when this guide looked at bridge operations earlier on, mainly holds bridge definitions, and it is what you use to move filters and basic settings around. A config backup ZIP holds the Bridges and entityMappings inside backup.json, and it says so explicitly: it contains no identity. Its archive does carry the per-bridge plugin enabled / disabled flags that exist at the time, but no bridge icon — and after you restore one, you are recommissioning every bridge in it, on purpose.
A full backup ZIP asks for the identity to be included, but it only packs one in for a bridge where that bridge's identity directory actually exists on disk, and it adds the matching bridge icon where it can. On the plugin side you still get nothing more than the enabled / disabled flags. includesIdentity: true in the archive metadata describes the request as a whole — it is not proof that every single bridge inside was packed with its identity intact. A stored snapshot, manual or automatic, uses that same conditional identity-and-icon scope, and retention quietly deletes the older files for you. After any restore, compare identitiesRestored against the number of bridges you actually selected before you decide the identities came through whole.
flowchart TD A["Four kinds of export"] --> B["Bridge export JSON:
bridge definitions only"] A --> C["Config backup ZIP:
Bridges + entityMappings,
no identity, explicitly"] A --> D["Full backup ZIP:
+ identity where the
directory exists,
+ matching bridge icon"] A --> E["Stored snapshot:
same conditional scope
as the full backup"] B --> Z["Never packed by any of
the four: plugin packages,
installed-plugins.json,
per-plugin config and secrets,
Lock Credentials, device-images,
app settings"] C --> Z D --> Z E --> Z
The code behind these archives simply does not pack every storage asset: installed plugin packages, installed-plugins.json, per-plugin config, storage and secrets, Lock Credentials and device-images all sit outside the backup scope. App-level settings such as Basic Auth, Auto Recovery and the backup preference itself do not appear in the backup.json restore scope either. Keep a separate note of the non-secret settings, keep a trusted original of every asset that is excluded, and put actual secrets nowhere but a dedicated secret manager and an encrypted backup — never in your own operations notes.
What each release channel actually promises
| Capability | Release channel | Product maturity | Controller support |
|---|---|---|---|
| Backup, preview, restore and snapshots | Stable 2.0.55 | A Stable operations capability — rehearse the destructive parts before you need them | A correct identity can keep the fabric, but a controller actually reconnecting still depends on your network |
| Auto Recovery | Stable 2.0.55 | A Stable setting | Retries failed bridges only; it never touches a bridge that is already running |
| Update Checker | Stable 2.0.55 | An informational feature | Shows different update guidance for add-on, Docker and npm installs |
| Restoring Server Mode and Camera / Security state | Backed up within Stable | Experimental-in-Stable | Restoring the data back is not the same claim as full controller support for it |
Update Checker compares your current version against the latest release information and shows the release notes and the runtime environment it detected. It does not install anything for you. The actual update method depends on whether you run the add-on, Docker or npm, so “a new version exists” is not the same sentence as “we have upgraded safely” — read the release notes, take a backup, confirm which image or package you can roll back to, and only then update inside a maintenance window.
Auto Backup is on by default with retention set to 5, and it attempts one automatic snapshot at graceful shutdown rather than on any periodic schedule. Auto Recovery is on by default with an interval of 60 seconds, and the UI lets you set anywhere from 10 to 3600. It restarts failed bridges periodically and also triggers a recovery pass after Home Assistant reconnects; a bridge that is already running is left alone. If the actual root cause is a port conflict, a bad setting or not enough memory, Auto Recovery only produces a growing pile of retry records — it will not repair the configuration underneath them.
Auto Recovery is a paramedic who only responds to a bridge that has already collapsed. It will not walk over to one that is standing there fine and start treating it, and if the collapse was a broken bone rather than exhaustion — a bad setting, not a transient hiccup — retrying the same treatment on a schedule does not set the bone.
Seven moves, in an order that keeps you from finding out the hard way
-
Step 1
Build a backup inventory
In Settings → Backup & Restore, create a manual snapshot, and download both a config backup and a full backup. Record the version, the date, the number of bridges, the number of mappings, and whether identity and icons were included. Never copy an identity or a secret value into that inventory — note that it exists, not what it is.
-
Step 2
Move it off the host, and verify it
Copy the archive to encrypted offline storage, compute a checksum, and confirm the ZIP actually opens and that you can read its README and metadata. Do not unpack the identity into a shared folder to test it, and clear any temporary files once you are done.
-
Step 3
Create a rollback point before any change
Before an upgrade, a plugin install, a migration or a restore, take one more full snapshot of the current state. Record the image or package version you are on, the storage mount, and the non-secret settings you would need to rebuild — and confirm the old image is still available to you.
-
Step 4
Run Restore Preview first
Upload a trusted archive and check its version, its creation time, whether it includes identity, whether each bridge already exists, and the mapping count. Select only the bridges you actually want restored. Leave overwrite off as it comes by default, and set include mappings and restore identity to match your plan — not the other way around.
-
Step 5
Stop the conflicting source, then restore
If you are migrating, stop the old instance first, so exactly one active instance ever holds that identity at a time. Run Restore and read the restored, skipped and errors counts it reports. If the UI asks for a restart, use a graceful restart and wait for Home Assistant and the bridges to settle before you judge anything.
-
Step 6
Verify identity stability, layer by layer
Confirm the version, Home Assistant connected, bridges running, the device and fabric counts, failed entities, sessions and subscriptions, and the mDNS interface. Then test one low-risk endpoint in both directions — controller to Home Assistant, and Home Assistant to controller. “The page loads” is not a complete verification on its own.
-
Step 7
If it fails, roll back — do not overwrite again
Stop the new instance, keep the redacted logs, and either restore the pre-change snapshot or start the old instance back up. Keep exactly one active instance at all times. If the archive reported errors, do not go on to overwrite more bridges on top of that — work out the version, the storage permissions and the missing files first, then retry only the items that actually failed.
identitiesRestored against the bridge count you selected — a mismatch means some identities did not survive.Moving hosts without losing commissioning, and keeping a small host alive
The safe migration order is fixed: full backup, then verify the archive, then stop the old instance, then restore the new instance onto the same storage layout with the complete identity, then start exactly one new instance, then verify. Only when the bridge ID, the identity directory and the endpoint identity all travel together do you have a real chance of controllers not needing to commission again. Copying just the bridge config, or letting the storage volume come up as an empty directory on the new host, will not get you there.
sequenceDiagram participant O as Old host participant Y as You participant N as New host Y->>O: take a full backup Y->>O: verify the archive, off-host Y->>O: stop the old instance Y->>N: restore the full backup,
same storage layout Y->>N: start exactly one instance Y->>N: verify identity, fabric, mDNS
Stable identity and persistent entity identity cut the chance that endpoints get renumbered after a restart, but a large change to the filters, the mappings, or a plugin's device set can still change what a controller sees on the other side. Do not rename things, rewrite filters, switch on Server Mode or jump several versions in the same move — prove the identity survived unchanged first, then change exactly one item at a time.
Keep a separate note of everything the archive leaves out
Track the non-secret settings on your own: whether Basic Auth comes from the environment or from stored settings, Auto Recovery's enabled state and interval, backup auto and retention, the mDNS start options, the base path and the log level. For an actual secret value, write down only that a secret manager supplies it — never the value. Bridge icons do travel in the full backup; device images do not, so keep the original files ready separately, and a local translation override has to be exported from the browser that holds it, since nothing else collects it for you.
When the host itself is the constraint
The official low-resource guidance for this version says Matter Hub loads the Matter cluster definitions, the Home Assistant registry and the V8 engine overhead at startup, and memory then grows with the endpoint count. When resources are tight: shrink the filters first, cut the endpoint count, turn off auto composed where you do not need it, move non-essential large add-ons elsewhere, and watch the heap and RSS trends in your metrics alongside any host-level out-of-memory signals.
Force Sync being skipped under heap pressure is a delivery van already loaded to the roof. Radioing the driver again and again to squeeze one more box in does not create space in the van — it just wastes the driver's time. The actual fix is fewer boxes, meaning fewer endpoints, or a bigger van, meaning more memory. Nothing about pressing the button harder changes how much room is left.
Force Sync can be skipped outright under heap pressure, and a process that restarts with no stack trace — where the last thing you see is Killed or a plain container exit — is the classic sign of an out-of-memory kill. On plain Docker or npm you can tune the Node heap by the official guidance; in the add-on, the entrypoint sets that dynamically, so there is no UI option to invent here. Swap is a buffer against a spike, not an answer to an endpoint count that keeps growing.
| Resource-pressure signal | Do first | Avoid |
|---|---|---|
| The heap sits near its limit for a long time | Shrink the entity or bridge set, check the plugins | Running Force Sync again and again |
| An exit code, or a host out-of-memory event | Check the host events, add usable RAM, or reduce the load | Only turning on debug logging, which adds more load on top |
| A large Home Assistant request times out | Check the Home Assistant load and the message timeout setting | Resetting the fabric straight away |
| A start-up spike from several bridges at once | Adjust the startup priority | Pressing Restart All over and over |
Three situations that end well, and one that is a last resort by definition
Routine updates: Update Checker only tells you a version exists. Read the release notes, take a full snapshot, keep the image you are currently on, update one environment, and verify it layer by layer as in the steps above. If a schema or controller regression shows up, stop the new version and roll back to the old version and the original snapshot — do not keep resetting things inside the broken environment hoping it settles.
Moving to a new host: prepare the new instance's networking, IPv6, mDNS and persistent storage first, but do not start it on the same identity while the old one is still up. Stop the old instance, then restore the full backup, keeping the bridge config and the network identity stable. Only once that is confirmed working do you close the rollback window on the old instance.
Lost storage: if you hold a verified full backup, restore it on the same Stable version. If all you have is a config backup, accept that you are recommissioning, and clear the old relationships on each controller before you pair again. Only when there is no usable backup at all do you move to a disaster reset — and treat rebuilding the controllers, Matter Hub and the automations as a project of its own, not an afternoon task.
A disaster reset is changing every lock on the house because one door is sticking, rather than calling someone out to fix that one lock. Every key that used to work — every controller that was already paired — stops working at once, not just the one that was giving trouble.
What you are seeing, what to look at, and the fix that does not make it worse
| Symptom | Check | Safe fix |
|---|---|---|
| Commissioning fails, or the bridge is never found | The bridge is running and not yet commissioned; phone and hub on the same network segment; IPv6, the mDNS bound interface, multicast, and the operational firewall | Go back to a plain single subnet first, correct the interface and firewall, and open the commissioning window again — do not run factory resets one after another |
| No Response after commissioning already succeeded | Home Assistant connected; whether the fabric still exists; sessions and subscriptions; whether mDNS is advertising the wrong interface; the controller hub itself | Repair the shared network first; if you need more, confirm autoForceSync, then restart or Force Sync that one bridge; if the other controllers are fine, look at that one hub and its support first |
| mDNS drops in and out, or shows duplicate records | AP multicast and IGMP; an mDNS reflector; multiple interfaces; any unclean power loss; the real fabric count | Bind the LAN interface, fix multicast, do a graceful restart and wait out the cache TTL; run a commissioning cleanup only on fabrics that really are surplus |
| A bridge shows Failed | The status reason; Home Assistant; the port; storage permissions; memory; the plugins | Fix the root cause, then let Auto Recovery or a single-bridge restart retry it; if the recovery history keeps failing anyway, turn the retries off and handle it by hand |
| Only some entities are marked failed | The failed reason; Home Assistant unavailable; the filter; the device class; the mapping and composed links | Fix the source or the mapping and restart that one bridge — do not reset the whole bridge over a handful of entities |
| Restore Preview looks fine, but Restore itself reports errors | The error on each bridge; exists and overwrite; the version; archive integrity and storage permissions | Stop overwriting anything else and restore the current snapshot instead; reproduce it on an isolated copy, and once you understand it, retry only the items that failed |
| After a migration, the controller finds two services | Whether the old host or old container is still active; the mDNS cache | Get back to a single active instance at once, stop the other end cleanly, and wait for or clear the controller cache — do not reset both ends together |
| A low-resource host restarts without warning | Metrics; host out-of-memory events; container exits; the endpoint count; large plugins | Cut the entity count, stagger bridge startup, and add RAM or set a suitable heap — do not tighten the Recovery interval and manufacture a restart loop |
| Last resort: a disaster reset | That the full backup provably will not restore, the identity is unusable, and network and controller problems are ruled out | Keep an archive of the current state and redacted logs, remove the old bridge on each controller one at a time, then run Factory Reset against the running Matter Hub bridge, cross-check the fabric and commissioning status afterward, and rebuild and commission again one bridge at a time; rebuild rooms and automations last |
flowchart TD A["A bridge or a controller
is not behaving"] --> B["Commissioning itself fails,
or the bridge is never found"] A --> C["It commissioned fine before,
and now shows No Response"] A --> D["The bridge itself shows Failed"] C --> E{"Is it one controller,
or every controller?"} E --> F["Just one: look at that
controller and its hub first"] E --> G["All of them: repair the shared
network, then Force Sync or
restart that one bridge"]
On this version, the very last row above has one extra wrinkle worth its own diagram: Factory Reset only actually runs against a bridge that is already Running. If you call it against a bridge that is Stopped or Failed, that bridge may simply be started — with no reset having run at all. Cross-check the fabric and commissioning status after the action either way, because a bridge that just started can look superficially like one that was reset.
stateDiagram-v2 direction LR state "Stopped" as Stopped state "Failed" as Failed state "Running" as Running state "Reset actually runs" as Done Stopped --> Running: reset call just starts it Failed --> Running: reset call just starts it Running --> Done: reset runs for real
The questions that come back around
Can a config backup keep my existing controller commissioning?
Does a full backup include the plugins, Lock Credentials, device images and every setting?
Does Auto Recovery ever restart a healthy bridge?
Does Update Checker upgrade Matter Hub automatically?
During a move, can I run the old and the new host together to test?
When is a factory reset actually the right move?
Where to go from here
You now have bridges that commission, filters that behave, and a backup you have actually tested.
This guide started with what a Matter bridge is and ends with keeping one alive: a backup that restores instead of just downloading, a migration order that does not force every controller to recommission, and a troubleshooting table that reaches for a reset only once everything else is ruled out. If a step here does not match what you see on screen, the earlier parts on bridges, filtering, mapping and commissioning are the ones to revisit first — most backup and No Response problems trace back to a setting from much earlier in the guide.
Revisit the full guidePart 13 of the Home Assistant Matter Hub Complete Guide series on the Apporo blog.
Adapted from the Home Assistant Matter Hub Complete Guide, produced by WoowTech and released under CC BY 4.0. This adaptation is published by Apporo under the same licence.
Light · Air · Water · Control · apporo