Skip to content

EMQX Enterprise 5.10.5

Choose a tag to compare

@zmstone zmstone released this 17 Sep 11:21
· 8773 commits to master since this release

Download

Ubuntu / Debian

OS Arch Package Tarball
ubuntu24.04 amd64 .deb (sha256) .tar.gz (sha256)
ubuntu24.04 arm64 .deb (sha256) .tar.gz (sha256)
ubuntu22.04 amd64 .deb (sha256) .tar.gz (sha256)
ubuntu22.04 arm64 .deb (sha256) .tar.gz (sha256)
ubuntu20.04 amd64 .deb (sha256) .tar.gz (sha256)
ubuntu20.04 arm64 .deb (sha256) .tar.gz (sha256)
debian13 amd64 .deb (sha256) .tar.gz (sha256)
debian13 arm64 .deb (sha256) .tar.gz (sha256)
debian12 amd64 .deb (sha256) .tar.gz (sha256)
debian12 arm64 .deb (sha256) .tar.gz (sha256)
debian11 amd64 .deb (sha256) .tar.gz (sha256)
debian11 arm64 .deb (sha256) .tar.gz (sha256)

RHEL / Rocky / Amazon Linux

OS Arch Package Tarball
el10 amd64 .rpm (sha256) .tar.gz (sha256)
el10 arm64 .rpm (sha256) .tar.gz (sha256)
el9 amd64 .rpm (sha256) .tar.gz (sha256)
el9 arm64 .rpm (sha256) .tar.gz (sha256)
el8 amd64 .rpm (sha256) .tar.gz (sha256)
el8 arm64 .rpm (sha256) .tar.gz (sha256)
amzn2 amd64 .rpm (sha256) .tar.gz (sha256)
amzn2 arm64 .rpm (sha256) .tar.gz (sha256)
amzn2023 amd64 .rpm (sha256) .tar.gz (sha256)
amzn2023 arm64 .rpm (sha256) .tar.gz (sha256)
el7 amd64 .rpm (sha256) .tar.gz (sha256)

macOS

OS Arch Package
macos14 arm64 .zip (sha256)
macos15 arm64 .zip (sha256)
macos26 arm64 .zip (sha256)

Plugins

Plugin Version Package
emqx_backup_sync 0.1.0 .tar.gz (sha256)
emqx_offline_messages 2.0.1 .tar.gz (sha256)
emqx_relup 1.0.2 .tar.gz (sha256)
emqx_sync_request 0.1.0 .tar.gz (sha256)

Breaking Changes

Access Control

  • #17864 The Dashboard user and API-key endpoints now reject scope lists that mix privilege scopes (system, user_management, api_key_management, sso_management) with other scopes. Each of the four privilege scopes is administrator-equivalent in effect, so combining them with a restricted scope list cannot meaningfully restrict the account. Use either a privilege-only scope list or a non-privilege-only scope list, depending on whether the account should have administrator-equivalent capability. Pre-existing records with a mixed scope set continue to function until the next update; the next update must split the list to succeed.

Data Integration

  • #18465 Fixed handling of templated INSERT SQL statements in the ClickHouse, TDengine, SQL Server, and MySQL bridges (when batch insert is enabled).

    Previously, rendering SQL templates could often produce malformed SQL due to syntax errors in the manually entered template itself and due to interpolation issues.

    Now, SQL statements are fully parsed when an action is created, and invalid SQL is rejected. During rendering, correct escaping is enforced. To provide consistent and predictable behavior, we limit the SQL features that can be used. Most notably, we reject comments in SQL statements. However, we support a large subset of syntax features: constant values, strings and string interpolation, arithmetic, functions, conditions, and conditional operators.

    MySQL also supports ON DUPLICATE KEY UPDATE, ClickHouse supports FORMAT Values and FORMAT JSONCompactEachRow, and TDengine supports INSERT ... USING ... TAGS and table identifier interpolation.

    To provide consistent rendering for MySQL templates, the MySQL bridge unconditionally disables ANSI_QUOTES and NO_BACKSLASH_ESCAPES modes for all connections, and treats the statements accordingly.

    The ClickHouse bridge now infers the batch value separator from the SQL template and ignores the configured batch_value_separator value.

Deployment

  • #17593 Added --force flag to emqx ctl relup upgrade. By default, the upgrade now refuses to proceed if data/patches/ contains any *.beam hot-patch files (which would shadow modules from the upgrade target). Pass --force to keep the patches and proceed anyway.

CLI

  • #18823 Fixed a misspelled field name in the emqx ctl listeners output.

    The command printed the listener's enabled flag as enbale. It now prints enable. Scripts that parse this output must be updated to match the corrected name.

Enhancements

Plugins

  • #17449 Added the EMQX Backup Sync plugin to periodically synchronize selected configuration from a primary cluster to a secondary cluster by using the Data Backup APIs. The plugin supports configurable TLS options for HTTPS calls to the primary cluster.

  • #17887 Added the emqx_sync_request plugin for synchronous MQTT request/response flows through the EMQX REST API. It also provides node-local CLI diagnostics for request counters and current pending state.

Performance

  • #17964 No longer spawn buffer worker pools for authenticator and authorizer resources. These resources always query through simple_sync_query, which bypasses the buffer workers, so the previously spawned workers were idle and never used.

  • #18138 #18186 Improved deep-page queries in the subscriptions HTTP API by accumulating in-memory subscription rows on each target node, avoiding one RPC per pagination batch.

Data Integration

  • #17472 Reduced the overhead of IoTDB REST API connector health checks by using a bounded version query instead of listing all databases on each check.

  • #17948 Added AWS IAM role credential support to DynamoDB connectors.

    When both the access key ID and secret access key are omitted, EMQX obtains temporary credentials from an ECS task role or EC2 instance metadata and refreshes them before they expire.

  • #18917 Added IPv6 support to the Kafka, Confluent and Azure Event Hubs connectors.

    • bootstrap_hosts accepts bracketed IPv6 addresses, for example [::1]:9092 or [fd00::5]:9092,host2:9093.
    • The connectors can reach hostnames that resolve only to IPv6 addresses, and brokers that advertise IPv6 addresses.
    • The new socket_opts.ip_family option selects the IP address family. With the default auto, a hostname is tried over IPv4 first and then over IPv6. Set it to ipv6 to connect over IPv6 only, or to ipv4 to connect over IPv4 only.

    The upgraded Kafka client library also fixes a sync produce timeout. It could happen when SASL re-authentication ran while requests were still pending.

Deployment

  • #18031 Added Enterprise Linux 10 (EL10) packages, for Red Hat Enterprise Linux 10, Rocky Linux 10, and compatible distributions.

  • #18118 Start releasing macOS 26 (Tahoe) packages.

Bug Fixes

Core MQTT Functionalities

  • #17729 Fixed a transient "address already in use" error that could occur when updating the options of a WS or WSS listener (for example when rotating TLS certificates). Updating such a listener rebinds its port, and the operating system may not have released the old socket yet; EMQX now retries the rebind briefly instead of failing the update.

  • #17522 Periodically purge stale entries from the global session registry. Previously, when a session's owner process died without a clean unregister (for example, after a brief network split that prevented the unregister from replicating, or when one core's consensus check timed out during the down-event cleanup), the registry row could remain forever if the same clientid never reconnected. A new throttled background sweep on each core node now removes such rows. The sweep is bounded to at most 500 registry rows per second per node and runs no more often than once every 10 minutes, so it does not measurably affect broker throughput even on registries holding millions of sessions.

  • #17573 Reduced MQTT v5 user-property parsing cost from quadratic to linear.

    Previously a CONNECT, PUBLISH or SUBSCRIBE packet carrying many user-properties caused super-linear scheduler time on the owning connection process, because each parsed property was appended to the end of the accumulated list. Parsing now scales linearly with the number of entries while preserving their wire order.

  • #18356 MQTT connections are now refused until node startup completes, so listeners no longer serve traffic before authentication, authorization, and plugin security hooks are active. The GET /status endpoint reports HTTP 503 during this startup window, so load balancers can route around the node until it is ready.

  • #18538 Fixed the multi-tenancy client list not following a persistent session that reconnects under a different namespace.

    Previously, when a client resumed an existing session (clean_start=false) after its namespace changed, GET /api/v5/mt/ns/{ns}/client_list kept listing the client under the old namespace, and the new namespace's list did not include it. The client list and the per-namespace client count now always reflect the namespace the client connected with.

  • #18623 Stop MQTT listeners before stopping applications during node shutdown.

    Previously, listeners kept accepting and processing client traffic while the applications behind the publish path were already stopped. Publishing clients could then trigger a burst of hook_callback_exception errors in the log, for example from the rule engine, until the listeners stopped a few seconds later. Listeners now stop first, so no client traffic is processed during application shutdown.

    The node now also reports itself as not running in GET /status as soon as shutdown begins, so load balancers stop routing new connections to it.

  • #18676 A malformed packet received before the CONNECT packet now closes the connection instead of crashing the connection process.

Security Hardening

  • #18205 Strengthened validation of data backup archives during import so a backup file's contents are restored only into the table it is meant for.

  • #17451 #17553 Restricted backup file downloads so only dashboard administrators can download archives containing dashboard accounts or API key records, while API key callers can still download archives without those sensitive records.

  • #17531 Bumped jiffy app version to 1.1.4 to address a memory safety bug in JSON handling.

    Without this fix, EMQX may occasionally (very rare) segfault when processing JSON data.

  • #17652 Fixed a security issue where the Prometheus configuration API returned stored Authorization header values in push gateway headers. The API now redacts these values in responses.

  • #17857 Improved redaction of sensitive data in logs, traces, and API responses.

    Secrets are now consistently hidden in the affected places, including authentication and authorization backend query traces, JWT signing key material (carried inside jose_jwk records), HTTP connector request headers, Prometheus configuration API responses, and authenticator creation responses.

  • #18336 Read-only REST endpoints no longer return secrets in cleartext:

    • GET /listeners and GET /listeners/{id} now render the listener ssl_options.password as ******.
    • GET /exhooks and GET /exhooks/{name} now render the gRPC client ssl.password as ******.
    • Audit log entries for POST /license now record the request body as ******, so the license key does not appear in GET /audit results.

    Updating a listener or an exhook server with a body that contains the ****** placeholder keeps the stored secret unchanged.

    Upgraded HOCON to 0.46.3. This release renders sensitive values inside array-typed config fields as ****** and no longer prints sensitive field values in config validation error logs.

  • #18626 Upgrade QUIC stack to quicer-0.4.8 (msquic 2.5.7).

    Contains a security update for CVE-2026-32179

  • #18835 Stopped including the result of a successful configuration change in the cluster configuration sync debug logs.

    The result could carry compiled runtime state, such as the HTTP authenticator header templates, which held secrets that log redaction did not cover.

Access Control

  • #17645 Fixed an HTTP/1.1 protocol-conformance issue in the JWKS retrieval client used by JWT authentication. Earlier versions sent an empty TE: header value due to a long-standing default in Erlang/OTP's inets HTTP client (fixed upstream in inets 9.4.2 / OTP 28.1). Some identity providers (notably PingFederate) reject such requests with 503 or a TCP reset. EMQX now sends an explicit valid TE: trailers header on JWKS fetches.

  • #17643 Fixed an issue where the plain password hash algorithm accepted passwords that differed only by letter case during authentication.

  • #17651 Fixed an issue where creating an authenticator via POST /authentication returned the new authenticator config without redacting provider secrets (such as JWT HMAC secrets, HTTP Authorization headers, and request body passwords). The creation response now applies the same redaction as the list and get endpoints.

  • #18147 Hardened scope-based authorization for the dashboard and management API so that access-control checks are applied consistently across equivalent request paths.

  • #18200 Fixed an error when updating an API key that was created with the scope left blank.

    The API key create and update requests now accept a scope list that matches the role's implicit default (and the unset value) as equivalent to "no explicit scopes", so re-submitting the value returned by a read no longer fails. Such keys also keep their forward-compatible implicit scopes instead of a frozen list.

    Also fixed the default administrator user being created with an explicit scope list at startup. The default administrator now follows its role's implicit default scopes (shown as unset), so it automatically gains scopes introduced in future releases instead of being pinned to a frozen list. Existing default administrator records that carry an explicit list are updated to the implicit form at boot.

  • #18963 Fixed POST /api/v5/api_key returning HTTP 500 when the optional desc or enable field is omitted from the request body. The key is now created with an empty note and enabled by default. Request body fields that are not part of the API key schema (for example description instead of desc) are still ignored by request validation.

  • #18997 Fixed an error when saving a Dashboard user whose scopes match the default set of its role.

    Such a user could not be edited at all, and switching its role between administrator and viewer failed with Privilege scopes cannot be combined with other scopes or Non-administrator users cannot hold admin-only scopes. The user API now reads a scope list that matches the role default, and the value unset, as "no explicit scopes". A user saved this way follows its role default as that default changes, instead of keeping a fixed list.

Data Integration

  • #17538 Fixed Redis Sentinel resources sharing one global Sentinel manager. Multiple Redis Sentinel resources on the same node now keep independent Sentinel server lists and credentials, preventing one resource from connecting through another resource's Sentinel configuration.

  • #17627 Fixed PostgreSQL connector batch execution when prepared statements are disabled.

    Previously, concurrent batch actions using different SQL templates could interleave on the same PostgreSQL connection during raw SQL batch execution. This could cause protocol_violation or invalid_sql_statement_name errors.

  • #17701 Fixed a confusing badarith error from PostgreSQL actions when a batched SQL template returned rows, for example SELECT ....

    PostgreSQL action batching does not support row-returning SQL. EMQX now returns a clear unsupported SQL error instead of crashing the batch result handler.

  • #17954 Fix GreptimeDB async batches that could remain unflushed after health checks at low write rates.

  • #17414 Fixed an issue where the health check of an Azure Blob Storage Connector could timeout, or generate large bandwidth costs, if the storage account contained too many containers. Companion fix to #16935.

  • #17567 Upgraded Kafka client library brod from 4.5.4 to 4.5.5.

    Fixed Kafka consumer group join failures against Kafka 2.2.0, where the broker returns member_id_required and brod previously discarded the assigned member ID instead of using it for the retry.

  • #17597 Fixed a connection failure to MongoDB 8.0+ when authentication is required. The driver previously queried buildInfo before authentication to pick the auth mechanism; MongoDB 8.0 restricted that command to authenticated callers. The driver now skips the probe and uses SCRAM-SHA-1 directly, which all supported MongoDB versions accept.

  • #17605 Fixed Oracle action prepare/status checks to parse action SQL without executing it, and reject unsupported top-level DDL/DCL/TCL statements. Also improved support for text payloads over 4000 bytes when the payload placeholder is the last bind parameter.

  • #17624 Fixed an issue with GCP PubSub Consumer Source where, if a source was initially created with a service account lacking necessary permissions to create subscriptions for the configured topic, the Source would fail to become connected even after granting the permissions to the service account.

  • #17716 Added the option of allowing TLS Peer Verification for Confluent Producer Connectors.

  • #17720 Added the option of allowing TLS Peer Verification for GCP PubSub Producer/Consumer Connectors.

  • #17947 Fixed an issue where updating an HTTP connector could leave its action buffer workers blocked after the connector was recreated, causing messages to remain queued until the next retry interval.

  • #17961 Fixed an issue where a Kafka or Pulsar Connector would transition to a disconnected state on health check timeouts, potentially recreating its internal queue. Now, they transition to connecting.

  • #17994 Fixed Kafka producer action retry metrics. The retried, retried.success, and retried.failed counters on an action's metrics now reflect messages that the internal buffer re-sends after a broker reconnect, so an operator can tell whether retried messages ultimately succeeded or failed. Previously these counters stayed at 0 regardless of how many internal retries occurred. The success and failed counters are unaffected and are not double-counted.

  • #18082 Kafka producer: max_linger_time is honored again for memory-mode buffers (upgraded the Kafka client library wolff to 4.1.11). When less than a full batch is buffered, the producer waits up to max_linger_time for more messages before forming a produce request, reducing the produce request rate at high message rates; sending is immediate whenever a full batch's worth of data is available, and the publish path is never delayed. The default max_linger_time = 0 keeps the previous send-immediately behavior.

    This also applies to Azure Event Hubs and Confluent producer connectors, which use the same client library.

  • #18251 Fix GreptimeDB connectors that could fail to restart when a stale gRPC channel remained after a worker was force-stopped.

  • #18301 Elasticsearch action index and id values are now URL-encoded when composing the request path, so characters such as # or / in a templated value are treated as literal text within a single path segment instead of altering the request target. The JSON request body is not affected.

  • #18317 When reading GCP Connectors (GCP PubSub Producer/Consumer) that use JSON Service Account authentication via the HTTP API, now the values are redacted.

  • #18328 The Snowflake connector now applies its configured ssl options when connecting to Snowflake endpoints. Previously the connector ignored the ssl settings and never verified the server certificate.

  • #18762 Fixed an error reported by the TDengine action. When the action could not be found, the error named the connector's ID instead of the action's ID, which made the error read as if a valid connector ID was invalid.

  • #18846 Fixed SQL template rendering in data integrations. Doris batch inserts now use Doris-compatible syntax and escaping for text and binary values. MySQL templates now handle escaped dollar signs correctly.

  • #18920 Fixed GreptimeDB connectors that switched between connected and disconnected under heavy write load. The health check no longer waits behind pending writes, so it fails only when GreptimeDB does not respond.

Plugins

  • #17875 Fixed plugin management HTTP APIs to ignore stale unpacked plugin directories that are not present in the cluster plugin config and are not running locally.

    Such stale packages no longer appear in plugin list/detail/config/schema responses, cannot be acted on by plugin operation APIs, and no longer block reinstalling the same package through the HTTP install API. Configured pre-installed plugins are still visible and continue to follow the documented pre-install workflow.

    EMQX now logs an error on startup and HTTP API access when a plugin package is unpacked but is neither enabled nor disabled in plugins.states.

  • #17710 Fixed noisy failed_to_get_plugin_config_from_cluster warning when installing plugins via CLI.

    The emqx ctl plugins install command now installs plugins in fresh_install mode (matching the HTTP API behavior), which skips the cluster config lookup for newly installed plugins, avoiding repeated config_not_found_on_node warnings on every node in the cluster.

    Added --cluster flag to emqx ctl plugins install for cluster-wide installation. When specified, the plugin package is distributed to and installed on all running nodes in a single command.

  • #17934 Fixed plugin package installation loading code before validating the package's application declarations, configuration schema, and default configuration.

  • #18334 Fixed an issue where a plugin could fail to start after a node restart.

    A plugin that declares the emqx_plugins application as a dependency made node boot 10 seconds slower and was left enabled but not running after every restart. EMQX now ignores this dependency declaration and logs a warning. Bundled plugins no longer declare it. When a plugin fails to start within the start timeout, the error log now lists the declared applications that were not running.

  • #18338 Start plugins after all EMQX applications have started. A plugin may now declare any EMQX application in its applications list. Previously, a plugin that declared an application which starts late in the boot sequence (for example emqx_management) failed to start after a node restart.

  • #18469 The hot-upgrade (relup) plugin now validates the target version string and checks upgrade-path compatibility before it modifies any files. An incompatible or malformed upgrade package is rejected without deleting or overwriting the installed release.

Gateway

  • #17885 Fixed an issue where the LwM2M gateway could include sensitive REGISTER query fields such as password, secret, private_key, and access_token in registration/update MQTT reports.

  • #18044 Redact sensitive fields from structured CoAP packet debug logs, including LwM2M registration query parameters.

  • #17395 Fixed CoAP gateway observe notifications to honor the gateway.coap.notify_type setting and queue pending confirmable Observe notifications instead of silently dropping them while another notification is waiting for ACK.

  • #17427 Fixed the JT/T 808 gateway schema validation to allow empty or omitted registry and authentication URLs when allow_anonymous is set to true. Previously, the not_empty validator was applied to both fields regardless of the allow_anonymous setting, causing a 400 error when submitting an empty string for these URLs even though they are not used in anonymous mode.

  • #17581 Fixed the JT/T 808 gateway to use the phone number accepted during authentication as the connection identity, rejecting mismatched registration-code authentication attempts and subsequent uplink frames with a different phone number.

  • #18652 MQTT-SN now publishes configured Will messages when sleeping clients exceed their sleep duration and no longer publishes Will messages when clients disconnect normally.

  • #18825 Fixed CoAP Observe notifications not being retransmitted after a missing ACK, which could leave subsequent notifications blocked in the pending queue.

Clustering

  • #17770 Fixed configuration update commands (REST API and CLI) crashing with a function_clause crash report when the underlying cluster RPC layer aborted with an unexpected reason, for example {no_exists, cluster_rpc_mfa} when the cluster RPC tables were not yet available during node startup or recovery. Such failures are now returned to the caller as a structured error instead.

  • #18000 Fixed a startup crash-loop that could occur when a node using the community (single-node) license joins a cluster whose peers hold a clustering-capable license.

    Previously, if cluster membership was established before the peer's license was replicated to the joining node, the node would refuse to start with a SINGLE_NODE_LICENSE error and, under an automatic-restart supervisor, keep crash-looping. The node now waits a bounded grace period for the clustering license to sync before it starts. A cluster in which no node ever obtains a clustering license is still rejected after the grace period elapses.

  • #18013 Fixed an issue that could terminate a node while it joined a cluster whose persisted mqtt.max_packet_size differed from its local configuration. EMQX now skips listener refresh side effects before listener startup and creates the listeners from the synchronized configuration when the EMQX application starts.

  • #18861 Validate the options passed to emqx_router_tool:scan_missing_routes/1 and emqx_router_tool:reconcile_missing_routes/1.

    Invalid chunk or sleep_ms values were accepted silently and disabled the scan throttling, so the scan ran at full speed while the operator believed it was throttled. The tool now raises an error naming the offending option instead. Unknown option keys, such as a misspelled chunks, are rejected as well.

Observability

  • #17709 Fixed a JSON log formatter crash that could replace some debug-level log and trace events with a FORMATTER CRASH line.

    The crash happened when a log field held a tuple value (for example the result field of the authenticator_result and authentication_result authentication trace events) and the active formatter configuration did not carry a chars_limit setting. As a result, the events that show which authenticator accepted or rejected a connection were missing from the output. These events are now formatted correctly.

  • #17886 Exposed the publish quota-exceeded packet metric in Prometheus as emqx_packets_publish_quota_exceeded.

  • #18107 Fixed an issue where the dashboard metrics APIs (GET /api/v5/monitor_current and GET /api/v5/monitor) returned 500 INTERNAL_ERROR while a node was joining the cluster.

    While a joining node is restarting its applications, sampling its metrics fails; this failure is now tolerated: the APIs return the aggregate of the remaining reachable nodes and log a warning, instead of failing the whole request.

    Also fixed a spurious clear_monitor_metrics_rpc_errors warning that was logged on every successful DELETE /api/v5/monitor request.

  • #18694 Fixed an issue where querying the audit log could return an error for records created by SSO-authenticated users.

API

  • #18067 Fixed the file transfer files API (GET /api/v5/file_transfer/files) failing to list files whose names contain non-ASCII characters (e.g. Chinese).

  • #18385 Fixed an issue where submitting a configuration containing an invalid Unicode escape sequence through PUT /configs returned an internal error. Such requests now return a validation error that names the invalid escape.

  • #18816 #18820 Fixed PUT /api/v5/telemetry/status returning 500 INTERNAL_ERROR with an Erlang stack trace when the request body omits the enable field.

    The endpoint now returns 400 BAD_REQUEST with a validation message. The API documentation marks enable as required and no longer shows a default value for it, because the endpoint has never applied that default.