Skip to Content

Back up the identity, verify it works, then fix what breaks without resetting anything

the one you bookmark, not the one you read once
Matter Hub Guide · Part 13

Back up the identity, verify it works, then fix what breaks without resetting anything

Twelve parts back you built bridges, filtered and mapped entities, and commissioned them into Apple Home, Google Home and Amazon Alexa. None of that survives on its own. This last part covers the two things that decide whether all of it is still there next month: taking a backup that actually restores, and working through a failure — No Response, a stuck commissioning window, a host that will not stay up — without reaching for a reset before you have to. That order matters. A reset is at the very end of this part for a reason.

4
Different exports Matter Hub can produce — bridge export, config backup, full backup, snapshot — and none of the four packs plugins, secrets or Lock Credentials
5
How many automatic snapshots Auto Backup keeps by default before the oldest one is deleted
60s
Auto Recovery's default interval. The setting's own range in the UI runs from 10 to 3600 seconds
Why a file is not a restore

Downloading a file and being able to restore from it are two different claims

What can actually be recovered from a Matter bridge is: the bridge config, the entity mappings, the Matter identity and fabric credentials, the plugin enabled / disabled state of each bridge, and a handful of administrative assets. That list matters because of what is missing from it once the identity goes. If the identity is lost you can rebuild the bridge settings from a config backup, but every controller that already commissioned this bridge still has to commission it again — the identity is what a controller recognizes, not the settings behind it. And if you ever start that same identity running on two hosts at once, you get duplicate services and mDNS and session behavior that nobody can predict for you in advance.

Settings → Backup & Restore is where all of this lives: downloadable config backups and full backups, a restore preview, per-bridge selection, an overwrite switch, include mappings, restore identity, internal snapshots, auto backup and retention. A restore can overwrite your current settings and ask the application to restart, so the rule before every restore is the same one you will see repeated through this part: take one more backup of the current state first, so you have an explicit way back.

In plain terms

A config backup is a copy of the rental agreement. It tells you the address, the terms, who lives where. A full backup includes the agreement and a working copy of the front door key. Lose the key and the agreement is still perfectly true — you just cannot get back in without going through the whole business of getting a new key cut and re-registering it with everyone who checks it, which is exactly what recommissioning is.

A full backup is a secret, not a convenience file. Any archive that includes the identity carries real Matter key material and fabric credentials — the thing that keeps your existing commissioning alive. Never share it, never commit it to version control, never drop it into ordinary cloud storage. Keep it on encrypted, access-controlled offline storage where you can revoke access, and check its integrity on a schedule rather than assuming a file that opens is a file that restores.
What you are actually facingEvidence you need firstSafe fallback point
Just a mis-edited settingA bridge export or config backup, and a diff against itA selective restore that leaves the identity alone
The host or the storage volume failedA full backup, its version, and a checksum you trustStop the old instance, then restore in full
A controller shows No ResponseHealth, sessions, mDNS, and the logsFix the network or restart that one bridge — do not reset first
The identity is confirmed unusableProof the backup will not restore, plus a controller cleanup planA factory or disaster reset, as the last resort
flowchart TD
  A["What are you actually facing?"] --> Q1["Just a mis-edited setting"]
  A --> Q2["The host or the volume failed"]
  A --> Q3["A controller shows No Response"]
  A --> Q4["The identity is confirmed unusable"]
  Q1 --> F1["Selective restore,
identity left alone"] Q2 --> F2["Stop the old instance,
then restore in full"] Q3 --> F3["Fix the network or restart
that one bridge, no reset"] Q4 --> F4["Factory or disaster reset,
the last resort"]
Four questions, four dead endsFour branches leave the top box and none of them rejoin — each one runs straight down to exactly one fallback box and stops there. Whichever branch matches what you are looking at is the whole answer for that situation; it does not point you at a fifth option.
Four exports, four promises

Bridge export, config backup, full backup, snapshot — read the label before you trust it

A bridge export JSON, covered when this guide looked at bridge operations earlier on, mainly holds bridge definitions, and it is what you use to move filters and basic settings around. A config backup ZIP holds the Bridges and entityMappings inside backup.json, and it says so explicitly: it contains no identity. Its archive does carry the per-bridge plugin enabled / disabled flags that exist at the time, but no bridge icon — and after you restore one, you are recommissioning every bridge in it, on purpose.

A full backup ZIP asks for the identity to be included, but it only packs one in for a bridge where that bridge's identity directory actually exists on disk, and it adds the matching bridge icon where it can. On the plugin side you still get nothing more than the enabled / disabled flags. includesIdentity: true in the archive metadata describes the request as a whole — it is not proof that every single bridge inside was packed with its identity intact. A stored snapshot, manual or automatic, uses that same conditional identity-and-icon scope, and retention quietly deletes the older files for you. After any restore, compare identitiesRestored against the number of bridges you actually selected before you decide the identities came through whole.

flowchart TD
  A["Four kinds of export"] --> B["Bridge export JSON:
bridge definitions only"] A --> C["Config backup ZIP:
Bridges + entityMappings,
no identity, explicitly"] A --> D["Full backup ZIP:
+ identity where the
directory exists,
+ matching bridge icon"] A --> E["Stored snapshot:
same conditional scope
as the full backup"] B --> Z["Never packed by any of
the four: plugin packages,
installed-plugins.json,
per-plugin config and secrets,
Lock Credentials, device-images,
app settings"] C --> Z D --> Z E --> Z
Every path ends at the same exclusionAll four export boxes in the middle row send an arrow down into the same bottom box — whichever one you pick, that bottom box is excluded from it. A standalone mapping profile export is not collected into any of the four either; it is its own file on its own path.

The code behind these archives simply does not pack every storage asset: installed plugin packages, installed-plugins.json, per-plugin config, storage and secrets, Lock Credentials and device-images all sit outside the backup scope. App-level settings such as Basic Auth, Auto Recovery and the backup preference itself do not appear in the backup.json restore scope either. Keep a separate note of the non-secret settings, keep a trusted original of every asset that is excluded, and put actual secrets nowhere but a dedicated secret manager and an encrypted backup — never in your own operations notes.

What “mapping” means here. Entity mappings travel inside both the config backup and the full backup, and a restore lets you choose whether to apply them. A mapping profile is a separate, portable rule file on its own path. Neither one is a substitute for the Matter identity — they describe what a bridge does, not who a controller thinks it is talking to.

What each release channel actually promises

CapabilityRelease channelProduct maturityController support
Backup, preview, restore and snapshotsStable 2.0.55A Stable operations capability — rehearse the destructive parts before you need themA correct identity can keep the fabric, but a controller actually reconnecting still depends on your network
Auto RecoveryStable 2.0.55A Stable settingRetries failed bridges only; it never touches a bridge that is already running
Update CheckerStable 2.0.55An informational featureShows different update guidance for add-on, Docker and npm installs
Restoring Server Mode and Camera / Security stateBacked up within StableExperimental-in-StableRestoring the data back is not the same claim as full controller support for it

Update Checker compares your current version against the latest release information and shows the release notes and the runtime environment it detected. It does not install anything for you. The actual update method depends on whether you run the add-on, Docker or npm, so “a new version exists” is not the same sentence as “we have upgraded safely” — read the release notes, take a backup, confirm which image or package you can roll back to, and only then update inside a maintenance window.

Auto Backup is on by default with retention set to 5, and it attempts one automatic snapshot at graceful shutdown rather than on any periodic schedule. Auto Recovery is on by default with an interval of 60 seconds, and the UI lets you set anywhere from 10 to 3600. It restarts failed bridges periodically and also triggers a recovery pass after Home Assistant reconnects; a bridge that is already running is left alone. If the actual root cause is a port conflict, a bad setting or not enough memory, Auto Recovery only produces a growing pile of retry records — it will not repair the configuration underneath them.

In plain terms

Auto Recovery is a paramedic who only responds to a bridge that has already collapsed. It will not walk over to one that is standing there fine and start treating it, and if the collapse was a broken bone rather than exhaustion — a bad setting, not a transient hiccup — retrying the same treatment on a schedule does not set the bone.

Restoring across versions is not promised in either direction. The archive carries version metadata, but nothing in the sources guarantees moving freely between any two versions. The safest route is to prove a successful restore on the same Stable version first, and only then upgrade following the release notes. Before you cross versions at all, keep the original archive and an environment you can still boot on the old version.
Build, verify, restore

Seven moves, in an order that keeps you from finding out the hard way

  1. Step 1

    Build a backup inventory

    In Settings → Backup & Restore, create a manual snapshot, and download both a config backup and a full backup. Record the version, the date, the number of bridges, the number of mappings, and whether identity and icons were included. Never copy an identity or a secret value into that inventory — note that it exists, not what it is.

  2. Step 2

    Move it off the host, and verify it

    Copy the archive to encrypted offline storage, compute a checksum, and confirm the ZIP actually opens and that you can read its README and metadata. Do not unpack the identity into a shared folder to test it, and clear any temporary files once you are done.

  3. Step 3

    Create a rollback point before any change

    Before an upgrade, a plugin install, a migration or a restore, take one more full snapshot of the current state. Record the image or package version you are on, the storage mount, and the non-secret settings you would need to rebuild — and confirm the old image is still available to you.

  4. Step 4

    Run Restore Preview first

    Upload a trusted archive and check its version, its creation time, whether it includes identity, whether each bridge already exists, and the mapping count. Select only the bridges you actually want restored. Leave overwrite off as it comes by default, and set include mappings and restore identity to match your plan — not the other way around.

  5. Step 5

    Stop the conflicting source, then restore

    If you are migrating, stop the old instance first, so exactly one active instance ever holds that identity at a time. Run Restore and read the restored, skipped and errors counts it reports. If the UI asks for a restart, use a graceful restart and wait for Home Assistant and the bridges to settle before you judge anything.

  6. Step 6

    Verify identity stability, layer by layer

    Confirm the version, Home Assistant connected, bridges running, the device and fabric counts, failed entities, sessions and subscriptions, and the mDNS interface. Then test one low-risk endpoint in both directions — controller to Home Assistant, and Home Assistant to controller. “The page loads” is not a complete verification on its own.

  7. Step 7

    If it fails, roll back — do not overwrite again

    Stop the new instance, keep the redacted logs, and either restore the pre-change snapshot or start the old instance back up. Keep exactly one active instance at all times. If the archive reported errors, do not go on to overwrite more bridges on top of that — work out the version, the storage permissions and the missing files first, then retry only the items that actually failed.

After a restore, read identitiesRestored against the bridge count you selected — a mismatch means some identities did not survive.
After a restore, read the restored / skipped / errors counts before you touch anything else.
Migration and low resources

Moving hosts without losing commissioning, and keeping a small host alive

The safe migration order is fixed: full backup, then verify the archive, then stop the old instance, then restore the new instance onto the same storage layout with the complete identity, then start exactly one new instance, then verify. Only when the bridge ID, the identity directory and the endpoint identity all travel together do you have a real chance of controllers not needing to commission again. Copying just the bridge config, or letting the storage volume come up as an empty directory on the new host, will not get you there.

sequenceDiagram
  participant O as Old host
  participant Y as You
  participant N as New host
  Y->>O: take a full backup
  Y->>O: verify the archive, off-host
  Y->>O: stop the old instance
  Y->>N: restore the full backup,
same storage layout Y->>N: start exactly one instance Y->>N: verify identity, fabric, mDNS
Old host, then you, then new host, left to rightThe three participants sit in the order they are first named in the migration checklist itself, and every arrow runs top to bottom in one direction: nothing about the new host happens until the arrow that stops the old instance has already completed.

Stable identity and persistent entity identity cut the chance that endpoints get renumbered after a restart, but a large change to the filters, the mappings, or a plugin's device set can still change what a controller sees on the other side. Do not rename things, rewrite filters, switch on Server Mode or jump several versions in the same move — prove the identity survived unchanged first, then change exactly one item at a time.

Keep a separate note of everything the archive leaves out

Track the non-secret settings on your own: whether Basic Auth comes from the environment or from stored settings, Auto Recovery's enabled state and interval, backup auto and retention, the mDNS start options, the base path and the log level. For an actual secret value, write down only that a secret manager supplies it — never the value. Bridge icons do travel in the full backup; device images do not, so keep the original files ready separately, and a local translation override has to be exported from the browser that holds it, since nothing else collects it for you.

When the host itself is the constraint

The official low-resource guidance for this version says Matter Hub loads the Matter cluster definitions, the Home Assistant registry and the V8 engine overhead at startup, and memory then grows with the endpoint count. When resources are tight: shrink the filters first, cut the endpoint count, turn off auto composed where you do not need it, move non-essential large add-ons elsewhere, and watch the heap and RSS trends in your metrics alongside any host-level out-of-memory signals.

In plain terms

Force Sync being skipped under heap pressure is a delivery van already loaded to the roof. Radioing the driver again and again to squeeze one more box in does not create space in the van — it just wastes the driver's time. The actual fix is fewer boxes, meaning fewer endpoints, or a bigger van, meaning more memory. Nothing about pressing the button harder changes how much room is left.

Force Sync can be skipped outright under heap pressure, and a process that restarts with no stack trace — where the last thing you see is Killed or a plain container exit — is the classic sign of an out-of-memory kill. On plain Docker or npm you can tune the Node heap by the official guidance; in the add-on, the entrypoint sets that dynamically, so there is no UI option to invent here. Swap is a buffer against a spike, not an answer to an endpoint count that keeps growing.

Resource-pressure signalDo firstAvoid
The heap sits near its limit for a long timeShrink the entity or bridge set, check the pluginsRunning Force Sync again and again
An exit code, or a host out-of-memory eventCheck the host events, add usable RAM, or reduce the loadOnly turning on debug logging, which adds more load on top
A large Home Assistant request times outCheck the Home Assistant load and the message timeout settingResetting the fabric straight away
A start-up spike from several bridges at onceAdjust the startup priorityPressing Restart All over and over
Updates, moves, disasters

Three situations that end well, and one that is a last resort by definition

Routine updates: Update Checker only tells you a version exists. Read the release notes, take a full snapshot, keep the image you are currently on, update one environment, and verify it layer by layer as in the steps above. If a schema or controller regression shows up, stop the new version and roll back to the old version and the original snapshot — do not keep resetting things inside the broken environment hoping it settles.

Moving to a new host: prepare the new instance's networking, IPv6, mDNS and persistent storage first, but do not start it on the same identity while the old one is still up. Stop the old instance, then restore the full backup, keeping the bridge config and the network identity stable. Only once that is confirmed working do you close the rollback window on the old instance.

Lost storage: if you hold a verified full backup, restore it on the same Stable version. If all you have is a config backup, accept that you are recommissioning, and clear the old relationships on each controller before you pair again. Only when there is no usable backup at all do you move to a disaster reset — and treat rebuilding the controllers, Matter Hub and the automations as a project of its own, not an afternoon task.

In plain terms

A disaster reset is changing every lock on the house because one door is sticking, rather than calling someone out to fix that one lock. Every key that used to work — every controller that was already paired — stops working at once, not just the one that was giving trouble.

A disaster reset is not a troubleshooting shortcut. It destroys the existing fabric relationships and the controller-side tidying you have already done. Run it only when the identity is already lost or damaged, the full backup will not restore, and you have genuinely ruled out network and session problems first.
Symptom, check, fix

What you are seeing, what to look at, and the fix that does not make it worse

SymptomCheckSafe fix
Commissioning fails, or the bridge is never foundThe bridge is running and not yet commissioned; phone and hub on the same network segment; IPv6, the mDNS bound interface, multicast, and the operational firewallGo back to a plain single subnet first, correct the interface and firewall, and open the commissioning window again — do not run factory resets one after another
No Response after commissioning already succeededHome Assistant connected; whether the fabric still exists; sessions and subscriptions; whether mDNS is advertising the wrong interface; the controller hub itselfRepair the shared network first; if you need more, confirm autoForceSync, then restart or Force Sync that one bridge; if the other controllers are fine, look at that one hub and its support first
mDNS drops in and out, or shows duplicate recordsAP multicast and IGMP; an mDNS reflector; multiple interfaces; any unclean power loss; the real fabric countBind the LAN interface, fix multicast, do a graceful restart and wait out the cache TTL; run a commissioning cleanup only on fabrics that really are surplus
A bridge shows FailedThe status reason; Home Assistant; the port; storage permissions; memory; the pluginsFix the root cause, then let Auto Recovery or a single-bridge restart retry it; if the recovery history keeps failing anyway, turn the retries off and handle it by hand
Only some entities are marked failedThe failed reason; Home Assistant unavailable; the filter; the device class; the mapping and composed linksFix the source or the mapping and restart that one bridge — do not reset the whole bridge over a handful of entities
Restore Preview looks fine, but Restore itself reports errorsThe error on each bridge; exists and overwrite; the version; archive integrity and storage permissionsStop overwriting anything else and restore the current snapshot instead; reproduce it on an isolated copy, and once you understand it, retry only the items that failed
After a migration, the controller finds two servicesWhether the old host or old container is still active; the mDNS cacheGet back to a single active instance at once, stop the other end cleanly, and wait for or clear the controller cache — do not reset both ends together
A low-resource host restarts without warningMetrics; host out-of-memory events; container exits; the endpoint count; large pluginsCut the entity count, stagger bridge startup, and add RAM or set a suitable heap — do not tighten the Recovery interval and manufacture a restart loop
Last resort: a disaster resetThat the full backup provably will not restore, the identity is unusable, and network and controller problems are ruled outKeep an archive of the current state and redacted logs, remove the old bridge on each controller one at a time, then run Factory Reset against the running Matter Hub bridge, cross-check the fabric and commissioning status afterward, and rebuild and commission again one bridge at a time; rebuild rooms and automations last
flowchart TD
  A["A bridge or a controller
is not behaving"] --> B["Commissioning itself fails,
or the bridge is never found"] A --> C["It commissioned fine before,
and now shows No Response"] A --> D["The bridge itself shows Failed"] C --> E{"Is it one controller,
or every controller?"} E --> F["Just one: look at that
controller and its hub first"] E --> G["All of them: repair the shared
network, then Force Sync or
restart that one bridge"]
Only one of the three branches keeps goingThree symptom boxes leave the top box, and only the No Response branch feeds into a second decision; the commissioning-fails and Failed branches end where they are, each already carrying its own check-and-fix in the table above. The second decision itself ends in exactly two boxes, with no further branching.

On this version, the very last row above has one extra wrinkle worth its own diagram: Factory Reset only actually runs against a bridge that is already Running. If you call it against a bridge that is Stopped or Failed, that bridge may simply be started — with no reset having run at all. Cross-check the fabric and commissioning status after the action either way, because a bridge that just started can look superficially like one that was reset.

stateDiagram-v2
  direction LR
  state "Stopped" as Stopped
  state "Failed" as Failed
  state "Running" as Running
  state "Reset actually runs" as Done
  Stopped --> Running: reset call just starts it
  Failed --> Running: reset call just starts it
  Running --> Done: reset runs for real
Two states in, one state where the reset is realBoth Stopped and Failed carry an edge labeled “reset call just starts it” — the reset itself does not run along either of those edges. Running is the only state with an edge into the box where the reset actually happens, which is why you start the bridge first and confirm Running before you trust that a Factory Reset call did anything at all.
Pairing data belongs in the local UI, nowhere else. Whatever the commissioning screen shows you during a reset-and-rebuild cycle, it is shown there because that is where it is safe to be shown. Do not copy it into a document, a ticket, or a message to anyone helping you — describe the situation in words instead.
FAQ

The questions that come back around

Can a config backup keep my existing controller commissioning?
No. It explicitly contains no Matter identity, so every bridge in it needs to be commissioned again after a restore. To keep the existing fabric intact, you need a protected full backup or snapshot, not a config backup.
Does a full backup include the plugins, Lock Credentials, device images and every setting?
Not that whole scope. It includes the bridges, the entity mappings, the identity where one exists on disk, the matching icons, and the per-bridge plugin enabled / disabled flags. It does not include plugin packages, the plugin registry, per-plugin config, secrets, Lock Credentials, device-images, or general app settings.
Does Auto Recovery ever restart a healthy bridge?
No. Both the setting's own description and its actual behavior deal with failed bridges only. A bridge that is already running is left alone, on purpose.
Does Update Checker upgrade Matter Hub automatically?
No. It compares your version against the latest release and shows guidance for your specific deployment — add-on, Docker or npm. The backup, the update itself, and any rollback remain entirely yours to run.
During a move, can I run the old and the new host together to test?
Not on the same Matter identity. Stop the old instance before you start the new one, and if you need to roll back, stop the new one first — exactly one active instance at a time, always.
When is a factory reset actually the right move?
Only when the commissioning identity genuinely has to be cleared, when a full backup will not restore, or when you have already decided to commission again and you have a controller cleanup and rebuild plan in hand. On this version, confirm the bridge is Running first and cross-check the fabric and commissioning status after the action — a Stopped or Failed bridge may simply be started, with no reset actually run. An ordinary No Response, an mDNS problem, or a handful of failed entities does not need a reset first, or at all.
Next

Where to go from here

that is all thirteen parts

You now have bridges that commission, filters that behave, and a backup you have actually tested.

This guide started with what a Matter bridge is and ends with keeping one alive: a backup that restores instead of just downloading, a migration order that does not force every controller to recommission, and a troubleshooting table that reaches for a reset only once everything else is ruled out. If a step here does not match what you see on screen, the earlier parts on bridges, filtering, mapping and commissioning are the ones to revisit first — most backup and No Response problems trace back to a setting from much earlier in the guide.

Revisit the full guide

Part 13 of the Home Assistant Matter Hub Complete Guide series on the Apporo blog.

Adapted from the Home Assistant Matter Hub Complete Guide, produced by WoowTech and released under CC BY 4.0. This adaptation is published by Apporo under the same licence.

Light · Air · Water · Control · apporo

What you expose to the network, and what you let run inside the process