Skip to Content

Nothing is coming through, and the four records that say why

look it up, stop guessing
EMQX Guide · Part 7

Nothing is coming through, and the four records that say why

The most common complaint in a smart home is a device that is “set up correctly” and still receives nothing. MQTT gives you nothing to look at — no cable to wiggle, no light to check — so the temptation is to start changing things at random. You do not have to. EMQX writes down who connected, what they subscribed to, what it kept, and which setting actually won. This part turns those four records into a routine: the Clients and Subscriptions pages for connection facts, Monitoring for whether the broker is well, Retained Messages for where a message went, and the configuration layers for why your change did nothing.

4 layers
base.hocon, cluster.hocon, emqx.conf, environment variables — the last one wins
2 kinds
Statistics are a snapshot, Metrics accumulate. Reading one as the other is the classic mistake
4 tools
Slow Subscriptions, Topic Metrics, Delayed Publish and Alarms are Enterprise, not Open Source 5.8.9
Where guessing stops

A broker that works is not the same as a broker you can read

Up to here the series has been about making things happen: installing EMQX, connecting a client, writing an authentication rule, sending data on to somewhere else. Everything in this part is about reading what already happened. That is a different skill, and it is the one that saves your evening.

Three ordinary situations, and where each of them is answered.

The sensorA temperature sensor was flashed, joined the Wi-Fi and shows a solid blue light. Home Assistant shows nothing. Is it even connected to the broker? Monitoring → Clients answers that in one glance.
The topicIt is connected, and still nothing arrives. What did it actually subscribe to — not what you meant to type, what it sent? Monitoring → Subscriptions lists it, character for character.
The settingYou changed a setting in the Dashboard, restarted, and the old value came back. Nothing is broken. A higher configuration layer overwrote you, and the chain that decides this has four links.

None of those needs a terminal, and none of them needs guesswork. The routine below is the spine of this whole part: each decision lands on a page covered in one of the sections that follow.

flowchart TD
  A["A device is set up.
Nothing arrives"] --> B{"Is it listed in
Monitoring, Clients?"} B -->|"not at all"| B1["It never connected.
Check the listener,
the credential, the ACL"] B -->|"listed, offline"| B2["The session outlived
the connection. Read
Session Information"] B -->|"listed, connected"| C{"Is the topic listed
under Subscriptions?"} C -->|"no"| C1["It subscribed to a
different topic filter.
Compare them character
by character"] C -->|"yes"| D{"Does a test publish
from WebSocket Client
reach it?"} D -->|"it arrives"| D1["The broker is fine.
The publisher is
the problem"] D -->|"nothing"| D2["Check QoS and No Local
on that subscription row"]
The routine, in orderEach of the three diamonds is answered by one Dashboard page, and the branch that carries on down the ladder is always the rightmost one — every dead end hangs off to its left. The last diamond is the exception: both of its answers are conclusions. Running it in this order is what stops you investigating subscriptions on a device that never connected.
One boundary to learn before you go hunting. EMQX ships in an open-source edition and a commercial one, and four of the tools named in this part — Slow Subscriptions, Topic Metrics, Delayed Publish and Alarms — are EMQX Enterprise features. On Open Source 5.8.9 they are simply absent from the sidebar. If you cannot find one of them, your install is not broken and you have not missed a setting. Knowing where that line falls saves you an hour of looking for a menu that was never there.
Clients and sessions

Who is connected, and who only looks connected

The Clients page shows the clients connected right now — and the sessions that have not expired yet. Those are two different things, and mixing them up is the reason the list sometimes disagrees with what you can see with your own eyes.

A connection is one live MQTT channel between a client and EMQX. A session is the state EMQX keeps on that client's behalf: its subscriptions and its queued messages. When a connection drops, the session does not necessarily go with it. If the client asked for a persistent session and gave a session expiry interval, the session lives on for that long, waiting for the device to come back.

In plain terms

A connection is the guest standing at the hotel desk. A session is the booking with their name on it. The guest can step outside for ten minutes and the booking is still theirs — the room is not given to anyone else until the checkout time passes. So a name on the list does not prove somebody is in the building, and walking a guest to the door does not cancel their booking.

stateDiagram-v2
  direction LR
  state "Connected" as C
  state "Disconnected, session kept" as D
  state "Session expired and gone" as G
  [*] --> C: the client connects
  C --> D: link drops, or Kick Out
  D --> C: reconnects in time
  D --> G: expiry passes
  C --> G: no persistent session
  G --> [*]
Two clocks, read left to rightA dropped link and Kick Out share the single edge into the middle box, where the session is still sitting — which is why a kicked device reappears moments later along the arrow back. Only the two edges reaching the box on the right actually end the session.

Opening the list and finding one device

  1. Step 1

    Open the Clients list

    In the sidebar, pick Monitoring → Clients. It shows the currently connected clients by default, with columns for client ID, username, connection status, IP, heartbeat, session information and the time the connection completed.

  2. Step 2

    Find one device with a filter

    The search bar at the top does a fuzzy search on client ID or username. Expand the arrow on its right for the proper filter fields: connection status, a time range and a target IP. Select Column at the top of the list chooses which columns are shown; Refresh resets every filter and reloads from scratch.

  3. Step 3

    Look at one client in detail

    Click a client ID to open its connection detail page. The top right lets you refresh by hand and clear the session by hand. Further down is the list of topics this connection currently subscribes to.

  4. Step 4

    Kick a client out, if you have to

    Back in the list, tick a client and press Kick Out to break the connection. Read the next callout before you rely on this.

The IP address column puts the client's source IP and the port it came in on together in one field, which is handy when two devices sit behind the same router. The detail page then adds three things the list does not carry: the protocol version the connection uses — MQTT 3.1.1 or 5.0, for example — whether the session is to be cleaned up once the client goes offline, and, if it is offline already, the time it last went offline.

The top of the detail page splits into two panels: Connection Information on the left and Session Information on the right. The session panel is the one worth learning, because it holds six fields that explain most “it is connected but slow” complaints.

Session fieldWhat it tells you
Session expiry intervalHow long the session survives after the connection drops
Creation timeWhen this session was created — not necessarily when the current connection started
Number of subscriptionsHow many topics this client is holding open
Message queue lengthMessages waiting for a client that is not taking them fast enough
Inflight window lengthMessages sent and not yet acknowledged
QoS2 receive queue lengthThe QoS 2 handshake that is still mid-flight

Below those two panels come the traffic, message and packet statistics, which make it easy to read the in and out volume of a single client rather than the whole broker. The bottom of the page lists the topics it subscribes to right now, with Add Subscription to add a simple subscription by hand and Unsubscribe in the list to cancel one.

Kick Out is not a ban. It closes the connection by force, and nothing more. A client on a persistent session maps straight back onto the same session when it reconnects, which for a device set to reconnect automatically means it is back within seconds. If you need a client to stay away, that is a job for Authentication, the ACL or the Blacklist — not for kicking it out again every time you notice it.
Subscriptions and topics

What the device actually asked for

The Monitoring → Subscriptions page lists every subscription on the broker as one row per client ID and topic pair. This is where you find out that the device subscribed to home/sensor/temp while you have been publishing to home/sensors/temp all afternoon. There is no clever way to catch that; you read the two strings side by side.

Each row also carries the QoS and the new MQTT 5.0 subscription options. Three of them come up often enough to be worth knowing by name. Read them as three separate settings rather than three dials on the same thing: one is about your own messages being sent back to you, and the other two are about the RETAIN flag — one on the way out, one on subscribe.

Subscription optionWhat it doesValues
No Local Stops the server forwarding your own published messages back to you Set to 1, the server does not send you your own messages
Retain As Published Sets whether the RETAIN flag is kept when a message is forwarded. This has nothing to do with the RETAIN flag on a retained message itself On or off, per subscription
Retain Handling Decides when the server sends retained messages on subscribe 0 = send as soon as the subscription succeeds; 1 = send only when no earlier subscription exists; 2 = never send

The search bar on this page carries three filter fields by default — Node, Client ID and Topic — and expanding the arrow adds two more: QoS and the shared subscription name, Shared Name.

Two views of the same thing, and both matter. The Topics tab takes the topics currently subscribed on every node, removes the duplicates and lists what is left, with a fuzzy search over it. So Subscriptions counts per client and Topics counts per topic. If two devices subscribe to the same topic, Topics shows one row and Subscriptions shows two. When you want to know “is anybody listening to this at all”, use Topics. When you want to know “what is this device listening to”, use Subscriptions.

On a Topics row, Create Monitor in the Actions column takes you to Diagnose → Topic Metrics to set monitoring up for that topic. That is an Enterprise page — more on it later in this part.

When it is online but slow (Enterprise)

Being connected and being served promptly are different problems. When a device is clearly online but slow to receive messages, EMQX's Slow Subscriptions measures the latency from the moment a message reaches EMQX to the moment it finishes being sent. Enable it under Diagnose → Slow Subscriptions, and four settings become yours to set.

Stats Threshold — only latencies above this are recorded. The minimum you can set is 100ms.
Record limit — at most 1000 records are kept.
Retention — how long a record is kept before it is dropped, 300 seconds by default.
Calculation method — one of whole, internal or response, which decides which part of the journey is being timed.

The resulting list is sorted from the longest latency down, with five columns: Client ID, Topic, Duration, Node and Updated. Clicking a Client ID opens that client's detail page, which puts you back in the previous section with a specific device to look at.

Slow Subscriptions and Topic Metrics are EMQX Enterprise features. On Open Source 5.8.9 the Diagnose group in the sidebar does not show them at all. If your edition is Open Source, the substitute is the statistics columns on the Clients and Subscriptions pages — a rougher read, but the message queue length and inflight window length on a client's detail page will usually tell you whether a device is falling behind.
Is the broker well?

Two kinds of number, and four ways to read them

A broker that works is not the same as a broker that is healthy under load. To keep the message backbone of a house running you want to know how many connections and subscriptions there are right now, how much memory the node has left, and whether the rate of messages in and out has blown up. The Monitoring module gathers all of that in one place.

Before the pages themselves, the distinction that everything else rests on. EMQX splits its monitoring data into two kinds:

Statistics are integer gauges — a single number for one moment in time. How many subscriptions exist right now.
Metrics are integer counters that accumulate — the running total of bytes and messages since counting started.
In plain terms

Statistics is the fuel gauge. It tells you what is in the tank at this second, and it goes down as well as up. Metrics is the odometer. It only ever climbs, and a single reading tells you nothing at all — the number is only useful next to what it read an hour ago. People panic at a large Metrics figure the way nobody panics at a car with 120,000 miles on it.

Cluster Overview and the Nodes tab

  1. Step 1

    Open the Monitoring module

    In the sidebar, choose Monitoring → Cluster Overview. The first thing on the page is the cluster-level connection, subscription and topic counts.

  2. Step 2

    Switch the group to Nodes

    In the top half, switch to Nodes to expand every node in the cluster. A home add-on install is usually a single node, so the first entry here is your add-on itself.

  3. Step 3

    Read the resources and the version

    The node row shows the connection count, the version, uptime, the number of Erlang processes and memory and CPU usage. The Version column is the most direct place to confirm that the EMQX you are running is 5.8.9. Click the node name to open the node details, where the system paths and the log path are written out.

  4. Step 4

    Use the Metrics tab for the cumulative counters

    Switch to the Metrics tab and work through the four groups of counters one at a time. Always read them together with the time range you have selected — a counter without a window around it is not information.

A node drawn in gray has stopped. That is worth committing to memory, because the instinct on seeing flat numbers is to assume the chart is broken. Check the node status column first. The time series chart below the list can be set from the last 1 hour or 6 hours up to 7 days, which is where you watch connection counts and message volume trend rather than twitch. If you want the Dashboard to start counting again from zero, Reset Monitoring Data clears what has accumulated so far.

On the Metrics tab, EMQX's counters come in four dimensions:

Metrics groupWhat it counts
bytesBytes in and out
packetsMQTT packets of each kind, sent and received
messagesMessage counts, including QoS 0/1/2, received, sent, forwarded and dropped
eventsEvent counts, such as connections, sessions, authentication and access

The Statistics side has its own names, and these are the ones you will see again on the REST API and in Prometheus, so they are worth reading once slowly. Most of them are shown as a current value alongside an all-time maximum, which is what makes them useful for a long-term view.

connections.count / connections.max — the current connection count, and the all-time maximum.
sessions.count / sessions.max — the current session count, and its all-time maximum.
subscriptions.count — the total number of subscriptions right now, shared subscriptions included.
topics.count — the number of unique topics right now.
retained.count — the number of retained messages right now.
delayed.count — the number of delayed-publish messages right now.
Alarms and Delayed Publish are Enterprise. Expand Monitoring in the sidebar and Open Source 5.8.9 gives you Cluster Overview, Clients, Subscriptions and Retained Messages. Delayed Publish and Alarms sit alongside them in the commercial edition only. Their absence is a licensing boundary, not a setting you got wrong.

Four ways out of the Dashboard

The Dashboard is only one of the doors onto this data. If you would rather not open it — because you want a chart on the wall, or an alert on your phone — there are three other routes.

flowchart LR
  E["EMQX 5.8.9"] --> D["Dashboard
Monitoring pages"] E --> R["REST API
the same numbers, as JSON"] E --> S["System topics
starting with $SYS/"] E --> P["Prometheus"] P --> G["Grafana charts
Alertmanager alerts"]
The same numbers, four doorsPrometheus is the last of the four branches and the only one that continues past its own box; the other three end where they are drawn. Nothing here is exclusive — you can watch the Dashboard and scrape the same broker with Prometheus at the same time.

Prometheus is the one most people end up using, because it is what puts EMQX on the same chart as everything else you measure on the host, and what lets you draw a Grafana dashboard at home and have Alertmanager tell you when something is off. On the Dashboard's Management → Monitoring page, open the Integration tab and choose the Prometheus settings. EMQX supports two modes.

ModeHow it worksCoverage
Pull Prometheus polls EMQX's REST API on a schedule. You point the url in your Prometheus configuration at EMQX and the metrics start arriving Three endpoints: /api/v5/prometheus/stats for basic metrics and counters, /api/v5/prometheus/auth for access control, and /api/v5/prometheus/data_integration for rules, Connectors, Actions and Sinks
Push EMQX pushes the metrics to a Pushgateway, off by default, and Prometheus collects them from there Currently the basic metrics only

Pull is what most people use, and the official EMQX documentation recommends Pull too — precisely because the Pushgateway mode currently covers only the basic metrics and is not as complete as auth or data_integration.

The pull-mode API needs no authentication by default. That means anything that can reach the port can read your broker's metrics: connection counts, topic counts, the shape of your house's traffic. To turn Basic Auth on, create an API key in EMQX and put that key into your Prometheus configuration. Do this before you expose the port to anything wider than the machine Prometheus runs on.
Where the message went

Retained messages, and the two things that are not one

MQTT works on publish and subscribe, and the broker does not file messages away and keep them forever. A message goes out to whoever is subscribed at that instant, and then it is gone. The one exception is the retained message, and understanding it answers most of the “where is that value now” questions in a smart home.

When a client publishes a message with the RETAIN flag set, EMQX stores it. From then on, any client that newly subscribes to that topic receives that message immediately, without waiting for the next publish. By default it never expires, unless you delete it by hand.

In plain terms

An ordinary publish is a shout down the hallway. Whoever happens to be standing there hears it; whoever walks in a minute later hears nothing and has no way of knowing they missed anything. A retained message is a note taped to the door. It is still there tomorrow morning for whoever arrives next, and it stays there until somebody takes it down or tapes a new one over it.

sequenceDiagram
  participant P as A thermostat
  participant B as EMQX
  participant S as A dashboard, later
  P->>B: PUBLISH with RETAIN set
  B->>B: keep it as this topic's retained message
  S->>B: SUBSCRIBE, hours afterwards
  B->>S: the stored message arrives at once
  P->>B: PUBLISH an empty body with RETAIN
  B->>B: the stored message is cleared
Stored on the way past, and cleared the same wayBoth arrows leaving the thermostat are publishes to the same topic with RETAIN set; only the body differs. There is no delete command anywhere in this picture — publishing an empty message is how a retained message is normally removed.

The Retained Messages page

  1. Step 1

    Open the list

    In the sidebar, choose Monitoring → Retained Messages to see every retained message in the system right now — topic, QoS, publisher and publish time.

  2. Step 2

    Read one message's payload

    In that row's Actions column, press Show Payload and the content opens below. You can pick a format such as JSON or Hex, and Copy at the bottom right copies it out.

  3. Step 3

    Delete one, or all of them

    Press Delete on the row to remove one. Clear All clears the retained messages across the whole cluster — read the warning below before you use it.

  4. Step 4

    Try it yourself with the WebSocket Client

    In the sidebar, choose Diagnose → WebSocket Client, add a connection, and use subscribe and publish to check a topic and its retained behavior in a few seconds. This is also the fastest way to prove the broker is publishing at all when a device says it is not receiving.

At the top of that list, Refresh reloads it and Settings takes you to Management → MQTT Settings → Retainer, where the feature can be turned on or off and five parameters set.

SettingDefaultWhat it means
Storage TypeBuilt-in DatabaseThe storage backend
Storage Methodramram: memory only. disc: memory plus disk
Max Retained Messages00 means no limit. Past the limit, new messages replace old ones
Max Payload Size1MBAnything larger is treated as an ordinary message and not retained at all
Message Expire IntervalNever0 means it never expires. You can have retained messages deleted automatically after a number of hours
Clear All is broader than it looks, and three of the entries are not yours. By default EMQX keeps three retained messages on $SYS system topics — the node description, the version and the cluster node list, for example. Clearing everything clears the whole cluster's retained messages, including the state your devices depend on for their first reading after a restart. The usual, surgical way to clear one topic is to publish an empty message to that topic, which also overwrites the old one.
A payload over the limit is not an error. A message larger than Max Payload Size is not rejected — it is delivered as an ordinary message and simply not retained. Subscribers connected at that moment get it and nobody sees a failure. If a large payload seems to be “disappearing” for late subscribers, this is the first thing to check.

Topic Metrics: one topic, counted (Enterprise)

Diagnose → Topic Metrics counts message volume for one specific topic and nothing else. Add a monitor there, or use Create Monitor on a Topics row, and give it a topic name. In the list's Actions column, View opens the details broken down by QoS, Reset starts the count over and Delete removes the entry.

Topic Metrics takes a complete topic name, not a topic filter. The official documentation spells it out: the feature supports a single topic name only, and the wildcards + and # are not supported at present. Something written as a/+ cannot be turned into a metric — the form will reject it. To watch a family of topics, you add one definite topic name at a time.

Delayed Publish: the other thing that is not retain (Enterprise)

Delayed Publish is an EMQX extension to MQTT. When the topic a client publishes to starts with $delayed/, EMQX holds the message back for a while and then publishes it on the real topic. The format is:

$delayed/{DelayInterval}/{TopicName}
$delayed/15/x/y — published to x/y 15 seconds later.
$delayed/60/a/b — published to a/b one minute later.
$delayed/3600/$SYS/topic — delivered to $SYS/topic one hour later.

{DelayInterval} can go up to 4294967 seconds. If it cannot be parsed as an integer, EMQX drops the message, so a typo in that segment costs you the message rather than delaying it. To manage the feature, open Management → Delayed Publish in the sidebar, where you can enable or disable it and cap the number of delayed messages.

Retained and Delayed Publish are different tools, and both are easy to reach for by mistake. Retained means “remember the last message so every new subscriber gets it immediately.” Delayed Publish means “send this later.” They can be used together, but they answer different questions — and Delayed Publish is an Enterprise feature, so on Open Source 5.8.9 the sidebar entry is not there.
Which layer wins

Your change did nothing, and here is the chain that explains it

EMQX parameters are scattered across several places: the ones you adjust dynamically in the Dashboard, the ones tucked away in a config file, and the ones you override with an environment variable. In a home setup it is easy to write the same setting into two layers by accident and then find your change does nothing, with no obvious way to tell which layer won.

Start from the two kinds of directory. Static configuration lives in etc and is usually read-only — the main config file, emqx.conf, is here, and you mostly change it at deployment or upgrade time. Dynamic configuration lives in data/configs and is writable: changes you make from the Dashboard, the REST API or the CLI are saved into cluster.hocon.

Then the chain. Precedence runs lowest to highest as base.hocon — which only exists from 5.8.4 on — then cluster.hocon, then emqx.conf, then environment variables. The further along that chain, the higher it wins.

flowchart LR
  A["base.hocon
weakest"] --> B["cluster.hocon
Dashboard writes here"] B --> C["emqx.conf
the static file"] C --> D["EMQX_ variables
strongest"]
Four links, one directionEach arrow means “is overridden by”, so read it left to right and the last box wins every argument. The trap sits in the middle: the Dashboard writes to the second box, and both boxes to its right can quietly overrule it on the next restart.
In plain terms

Imagine a shop with a printed staff handbook, a memo from the manager, and a note taped to the till. All three say something about the opening hours. Everyone works from the note on the till, because it is the one in front of them. Correcting the handbook is not wrong, but nothing changes on the shop floor while the old note is still stuck to the till.

LayerWhere it livesWho writes to it
base.hoconPresent from 5.8.4 onwardsThe weakest layer; you rarely touch it
cluster.hocondata/configs/cluster.hoconEvery dynamic change from the Dashboard, the REST API and the CLI
emqx.confThe etc directory. In the add-on, /data/emqx/etc/emqx.confYou, by hand, at deployment or upgrade time. It takes precedence over cluster.hocon
Environment variablesThe add-on's env_vars optionYou. Highest precedence of all — it overrides every config file
Do not set the same item in two layers. If a dynamic change of yours went into cluster.hocon but the same item is also set in emqx.conf or in an environment variable, it can be changed back after a restart. The recommendation is plain: pick one layer per setting. If you want a value pinned so nothing can move it, put it in the highest layer — an environment variable — and leave the Dashboard alone for that item.

The half of this you can do without a restart

Everything so far reads like configuration is a file you edit and a service you restart. Half of it is not. The Dashboard's Management module is the way in to dynamic changes, and four areas live there: Cluster Settings, MQTT Settings, Logging and Monitoring. Those can all be applied while EMQX is running, with no restart.

Another way to hold the split: a Management page is a switch you flip on a machine that is already running, while emqx.conf and env_vars are edits to the instructions the machine reads when it starts. The first shows up the moment you save. The second waits for the next start.

To see them, open the sidebar and choose Management → Cluster Settings. That is where you adjust MQTT, Listener and Logging settings, and they take effect as soon as you save. There is no file to find and nothing to restart. What you save is still a dynamic change, so it lands in cluster.hocon — which is why a higher layer can overrule it later, exactly as the chain above says.

Which half am I in? If the thing you want is on a Management page, change it there and it applies to the running broker. If it is not, it is a static item, and that is when the rest of this section — emqx.conf, an environment variable, and a restart — becomes the route. The restarts described below belong to the add-on's env_vars, not to configuration in general.

HOCON, and the two forms of the same line

From EMQX 5.0 on, configuration is written in HOCON, a superset of JSON. That means you can use nested objects or flat paths, and they mean the same thing. The nested form opens a node { } block and sets name and cookie inside it; the flat form puts the whole path on one line:

node.name = "[email protected]" — the flat form, one line per setting.
Inside a node { } block the same setting is written name = "[email protected]", alongside cookie = "<your-cookie>".

Environment variables and the add-on's env_vars

Three rules convert a config-file path into an environment variable, and all three have to hold or nothing happens.

A config file separates levels with a dot, . — an environment variable uses a double underscore, __.
Variable names always start with EMQX_.
The value is parsed as HOCON, so complex types can be passed — and so special characters need care, as below.

In the Woow EMQX add-on this is exposed through the env_vars option in config.yaml. Each entry is a name and a value pair:

namevalueWhat it sets
EMQX_NODE__NAME"[email protected]"The node name. The double underscore stands for one level of the configuration, so this is node.name. That value is the one in the add-on's own example. The HOCON snippet earlier wrote the same setting as "[email protected]", and because an environment variable is the higher layer, the variable is the one that would win
EMQX_LISTENERS__TCP__DEFAULT__MAX_CONNECTIONS"1000000"The maximum connections on the default TCP listener — three levels deep, three double underscores
  1. Step 1

    Open the add-on's configuration

    Home Assistant → the EMQX add-on in the sidebar → Configuration. Fill in name and value under env_vars.

  2. Step 2

    Use a name the add-on will accept

    The add-on's config.yaml validates environment variable names against the regular expression ^EMQX_([A-Z0-9_])+$. Only names starting with EMQX_, in capitals, get through. A misspelled name shows up as a warning in the startup log rather than as a rejected form.

  3. Step 3

    Save and restart the add-on

    Nothing takes effect until you save and restart. This is the single most common reason an env_vars entry “does not work”.

  4. Step 4

    Read the value back in the Dashboard

    Change one thing, restart, and confirm the new value in the Dashboard before you change the next. If it did not take, you now know it is that one entry and not a combination.

If you do need the static file. The add-on's data directory holds /data/emqx/etc/emqx.conf. Do not edit it by hand unless you have to, because everything the Dashboard writes lands in cluster.hocon, and hand-editing the higher layer is how you end up with two answers to the same question.

Secrets, and the comment trap that eats your password

EMQX configuration has a Secret type for sensitive values such as passwords and tokens. Two habits go with it. First, whenever you write down what goes into an add-on field or a Dashboard field — in a note to yourself, in a document, in a message to somebody helping you — use a placeholder such as <your-password> in place of the real value, and never the actual secret.

Second, and this one bites people who did nothing wrong: in HOCON, # starts a comment. If a password contains a # — say MQtt#123 — and it is not quoted, the parser reads it as MQtt and throws #123 away as a comment. Your password is now four characters long and you have no idea why login fails. Wrap it in HOCON-level double quotes so the whole literal survives:

export EMQX_DASHBOARD__DEFAULT_PASSWORD='"MQtt#123"'

Note the shape of that: single quotes on the outside for the shell, double quotes on the inside for HOCON. Do the same for any value containing : or =. And do not reach for URL encoding to dodge the problem — %23 does not help here, because EMQX does not decode URL encoding in environment variables.

The password above is an example, not a password. MQtt#123 is the documentation's illustration of the comment trap. Do not use it, or anything shaped like it, on a broker you actually run. The same goes for every default credential EMQX ships with: change it before the broker is reachable by anything but you.
When it does not add up

The symptom you see, and the page that explains it

These are the situations that come up most across the four areas in this part, matched to what to do about them. Read the symptom column first; most of these look like faults and are not.

SymptomLikely causeWhat to do
A device is missing from Clients It never connected, or you are searching for the wrong string Fuzzy-search by client ID or username first and confirm it really did connect. If it is offline already, its session may still be listed, which makes it look online
The same client is back right after a Kick Out A client on a persistent session reconnects to that same session inside the session expiry interval To keep it out, use Authentication, the ACL or the Blacklist rather than kicking it by hand every time
Subscribed to a topic, and nothing arrives The subscription's QoS or No Local is set wrong, or the broker is not publishing at all Check that row on the Subscriptions page. Then subscribe to the same topic again with the tool in Diagnose → WebSocket Client, to find out which half of the problem you have
Slow Subscriptions or Topic Metrics is nowhere to be found Both are Enterprise features On Open Source 5.8.9 use the statistics columns on Clients and Subscriptions for a rough read instead
The numbers on the home page look stuck The Overview charts have a time range and an aggregation interval Widen the range and look again, or press Reset Monitoring Data to start accumulating from scratch
A node is drawn in gray That node has stopped Do not read the data as a missing entry. Check the node status column first, then go and find out why it stopped
You need to know which EMQX version you are running Nothing is wrong; you just need the number Go to Monitoring → Nodes. The Version column is the version currently running. This guide targets 5.8.9
Delayed Publish and Alarms have vanished from the sidebar Both are Enterprise features Open Source 5.8.9 does not offer them, and nothing is broken
You subscribed to a topic but no retained message arrived There is no retained message on that topic any more Check the Retained Messages page. Once it is deleted, it is not sent again — only a fresh publish with RETAIN puts one back
You want to clear a topic's retained message but cannot reach Delete You are not at the Dashboard, or the row is not where you expected The more common approach is to publish an empty message to that topic with RETAIN, which also overwrites the old one
Topic Metrics rejects what you type when you add one You gave it a topic filter with a wildcard The feature only supports a single topic at present, not + or #. Create it again with a complete topic name
A setting you changed is back to what it was after a restart Usually emqx.conf or an environment variable has the higher precedence Move the value you want up a layer — to an environment variable, say — and pin it there
env_vars has no effect The name is wrong, or the add-on was not restarted Check the name starts with EMQX_ and uses a double underscore __, then restart the add-on after saving. A misspelled name shows up as a warning in the startup log
A password containing # gets truncated HOCON treats # as the start of a comment Store the whole thing wrapped in an extra pair of double quotes, quotes included, rather than pasting the raw value
You cannot find the setting you want to change You are looking in the wrong kind of layer Work out first whether the Dashboard can change it dynamically or whether it is a static item in a config file. Use Management → Cluster Settings for the first, where the change applies with no restart; go to emqx.conf or an environment variable for the second
Questions people ask

The ones that come up again and again

What is the difference between Clients and Sessions?
The Clients page shows connections and sessions at the same time, which is why it can look inconsistent. A session does not necessarily disappear the instant its connection drops: if that client uses a persistent session with an expiry set, the session lives on until it expires. Session Information, in the top right of a client's detail page, is where the session's own fields are — the session expiry interval, the creation time, the number of subscriptions, the message queue length, the inflight window length and the QoS2 receive queue length.
Why does a client I kicked out connect again on its own?
Kick Out only breaks the connection; it does not forbid reconnection. A device set to reconnect automatically comes back within the session expiry interval it is allowed, and maps onto the same session it had before. To turn a client away again and again, look at blocking it with the Blacklist or at the ACL level. Kicking a device out by hand is a diagnostic step, not a control.
Are the Topics list and the Subscriptions list the same thing?
Not entirely. Subscriptions counts one row per client and topic pair; Topics lists topic names deduplicated across all nodes. If two devices subscribe to the same topic, Topics shows one row and Subscriptions shows two. Use Topics to answer “is anyone listening to this?” and Subscriptions to answer “what is this device listening to?”
Why can I not find slow subscription statistics and topic monitoring?
Slow Subscriptions and Topic Metrics ship with EMQX Enterprise only. The Diagnose group in the Open Source 5.8.9 sidebar does not show them. If your edition is Open Source, work from the statistics on the Clients page by hand — the message queue length and inflight window length on a client's detail page are the closest substitute for a latency reading.
What is the actual difference between Statistics and Metrics?
Statistics are integer gauges read in one shot: how many subscriptions there are right now, what the all-time maximum was. Metrics are counters that accumulate: how many bytes have been received in total, how many packets have gone out. The Dashboard keeps both under Monitoring. The practical consequence is that a Statistics number is meaningful on its own, and a Metrics number is only meaningful next to an earlier reading of the same counter.
Why can't I see Alarms or Delayed Publish in the Dashboard?
Because they are EMQX Enterprise features. Open Source 5.8.9 does not show them in the Monitoring sidebar. That is a licensing boundary, not a setting you got wrong, and no amount of restarting or reconfiguring will make the menu entries appear.
Should I pick Pull or Push for Prometheus?
Most people pick Pull, and the official EMQX documentation recommends Pull too, because the Pushgateway push mode currently covers only the basic metrics — it is not as complete as auth or data_integration. With Pull you point the url in your Prometheus configuration at EMQX and the metrics start arriving from /api/v5/prometheus/stats, /api/v5/prometheus/auth and /api/v5/prometheus/data_integration. Remember that the pull-mode API needs no authentication by default: create an API key in EMQX and put it into your Prometheus configuration to turn Basic Auth on.
How do I check one node's system resources?
Go to Monitoring → Nodes and click the node name to open its details page, which holds memory and CPU usage, the Erlang process count, the connection count and so on. On a single node at home, read that page as the health of your EMQX instance — there is nowhere else for the load to have gone.
Does a retained message expire on its own?
Not by default. It stays until you delete it, overwrite it with an empty message, or give it an expiry time in the Retainer settings under Message Expire Interval. You can also carry a different expiry in seconds on the PUBLISH packet itself, and the value on the PUBLISH wins over the Retainer setting.
Why do subscribers get the old value after I publish the retained message again?
Usually because that retained message has already been cleared, or because the new publish went out without the RETAIN flag — an ordinary publish reaches current subscribers and changes nothing that is stored. Publish a new message to the same topic with RETAIN and the retained message overwrites the old value in the same fields on the same topic.
Are Delayed Publish and retained the same thing?
No. Retained means “remember the last message so every new subscriber gets it immediately”; Delayed Publish means “send it later”, using the $delayed/ prefix. You can use both together, but they play different parts. Delayed Publish is an Enterprise feature, so on Open Source 5.8.9 the choice does not arise.
Why can't Topic Metrics use a/+?
The official documentation spells it out: Topic Metrics supports a single topic name only, which means the + and # wildcards are not supported at present. To monitor a topic, use one definite topic name. This is a limit of that one diagnostic feature, not of MQTT topic filters generally — ordinary subscriptions take wildcards as they always have.
Where do the settings I change in the Dashboard get stored?
Dynamic changes from the Dashboard, the REST API and the CLI are written into data/configs/cluster.hocon, at cluster level. If you do not want a value changed there, set it in a higher layer instead — emqx.conf or an environment variable.
Which wins, emqx.conf or cluster.hocon?
emqx.conf takes precedence over cluster.hocon. Even so, do not set the same parameter in both, or you will lose track of the final value after a restart. Pick one layer per setting and stay in it.
What if the same parameter is set in both an environment variable and a config file?
The environment variable wins. It overrides the config file. Its value is parsed as HOCON, though, so wrap it in HOCON double quotes when it contains a special character such as #, : or = — otherwise the parser will read the value as something shorter than you typed.
Which variable names does the add-on's env_vars accept?
The add-on's config.yaml validates environment variable names against the regular expression ^EMQX_([A-Z0-9_])+$, so you can only enter names that start with EMQX_. Levels of the configuration are separated by a double underscore __, so node.name becomes EMQX_NODE__NAME. Save and restart the add-on before you expect anything to change.
Next

Where to go from here

read it, do not guess it

You can now answer “is it connected, did it arrive, which setting won” without touching anything.

Part 8 goes one level below the Dashboard: the log files and what the levels actually mean, the REST API for the numbers you want on your own terms, ngrok for reaching the broker from outside, and the backup and recovery routine you want in place before you need it.

Open the full guide

Part 7 of the WoowTech EMQX Complete Guide series on the Apporo blog.

Adapted from the WoowTech EMQX Complete Guide, produced by WoowTech and released under CC BY 4.0. This adaptation is published by Apporo under the same licence.

Light · Air · Water · Control · apporo

in EMQX
Messages from somebody else's broker, and Home Assistant on yours