Harper Pro 5.3.0 is the first stable release of the 5.3 line. These notes cover everything since the 5.2 line branched at v5.2.4, across 5.3.0-alpha.1 and beta.1–beta.4, written as the difference from 5.2.14: changes backported to the 5.2 patch train are listed once, under "Already shipped in 5.2.x". Harper (core) moves to v5.3.0; its notes are inlined below the divider.
Headline harper-pro work: replication correctness (whole transactions per origin, a log-order anchor for joining nodes, preserved origin versions, a capability registry), one jittered retry schedule with bounded subscription setup, cluster-wide record locks, resumable clones, a working update_node, and SSH deploy-key hardening.
Upgrade notes
- Upgrade every node before relying on the SSH key fixes or running the delete-echo repair tool. A cluster is protected by the SSH key name validation only once every node runs it (#910), and
repairDeleteEchoRunsmust run only after every node is on 5.3 (#895). add_ssh_keyandupdate_ssh_keyrefuse keys ssh cannot use (public keys, PuTTY or passphrase-protected keys, DSA/FIDO keys, certificates, damaged pastes) andhost/hostnamevalues that would break the shared ssh config, with a400naming the mistake. Scripts that relied on storing such keys will now fail at the API instead of at deploy time. (#918)- Each node rewrites its
<rootPath>/ssh/configonce at startup to add# BEGIN/END harper ssh key <name>lines around each key's block. Hand edits outside those lines are preserved from then on. (#938) - A persisted
isLeader: falsenow overrides a configured leader (HDB_LEADER_URL, CLI, orroutes[0]); 5.2 ignored a persistedfalse. (#800) - SSH configs damaged by the pre-5.3 block-matching bug are not repaired automatically: run
get_ssh_key,delete_ssh_key, thenadd_ssh_keywith the returned sealed key and the correcthost/hostname. (#910)
Replication
Correctness
- Replicated transactions stay on one connection and one origin. One peer's frame can no longer mix into or split another peer's transaction; a frame spanning several origins is applied as one transaction per origin, a frame that ends without its end marker is aborted rather than partially committed, and a failed transaction is replayed a bounded number of times. (#907, with core #2800)
- A node joining mid-transaction no longer misses that transaction. The base copy published a timestamp-based resume point, so a transaction in flight at join time could be delivered by neither the copy nor the log tail. The copy now anchors on a log-order barrier. (#878)
- Origin record versions and transaction-log keys are preserved instead of landing under the receiving node's clock, which later audit reads and conflict resolution depend on (#812); the blob codec survives the trip instead of being re-encoded (#795).
- Peer table definitions are applied additively, so a partial
DB_SCHEMAsnapshot from a peer cannot destroy locally declared attributes. (#750) - Protocol capability registry: capabilities are declared in one place, so nodes of different versions negotiate features explicitly. (#813)
- Delete-echo repair tool. The echo itself is fixed in core; for transaction logs that already hold echoed delete runs,
node dist/bin/repairDeleteEchoRuns.js <harper-root>reports on a stopped node,--applyrepairs and keeps backups,--restore <backup-dir>undoes. Run it only after every node is on 5.3. (#895)
Retries and liveness
- One jittered backoff schedule for every replication retry site, with admission control on subscription setup. A transient boot-time DNS failure could re-drive subscription setup once per event with no dedup or cap — about 1,400 log lines per second per node, ending in an out-of-memory kill. Setup now holds at most one armed attempt per peer and database (200 ms floor, jittered up to 30 s, reset on connect), reconnects use full jitter under the existing 500 ms–30 s bounds, recovery timers are owned and cancelled with their entry, and setup never runs on the main thread. (#800)
- An unanswered replicated operation fails when the peer connection closes or at a caller-supplied deadline, instead of waiting forever; each failed peer is logged at warn. (#915)
- A setup-recovery convergence timeout reports the replication state it was waiting on (#868); authorization denials are logged with the user and peer identities before the connection closes with 1008 (#894); worker exits record which link-status updates were not stamped and why (#843).
Cluster record locks
The operator-agreed home map is transported between nodes and successor-freshness barriers are enforced, completing the cluster half of core's table.lock(id). (#822, core issue #2542)
Nodes and clones
- A clone's sync wait is bounded by the data size the leader reports and resumes across restarts without redoing setup, so a large clone no longer times out on a fixed budget or starts over. (#661) The v4 clone source gate is held through a delayed reconnect (#839).
update_nodeworks. The documented operation was never registered. A request that only changesrevoked_certificatesorshardnow patches the node record locally without resetting replication topology. (#916)remove_noderevokes locally before notifying the peer, so a removed node cannot keep authenticating during the notification window. (#752)- Cloning no longer logs a misleading error when the leader has no env-secrets key, and failed leader requests include the leader's reason (#914). A failed SSH key clone is contained like a failed JWT one, and a failed leader request is retried (#918).
Security
- SSH key operations are confined to their own key.
get_ssh_key,update_ssh_keyanddelete_ssh_keydid not validate the key name, so a name such as../keys/.jwtPrivatecould read, overwrite or delete files outside the ssh directory, including the JWT signing key, and update/delete replicated that to every peer. All key operations now share one name rule, checked before anything replicates. Config blocks are matched exactly, so deletingrepono longer removesrepo-2's block. (#910) - SSH deploy keys are validated and normalized on add/update (indented or CRLF pastes now work), and the ssh config is written under one lock with rollback of a partly written append. (#918, #938)
- Forwarded operations no longer carry the requesting user's refresh token or log that user record. (#915)
Certificates, build and operations
add_certificatefinds a matching private key already on disk (for example an auto-generatedprivateKey<uuid>.pem) instead of failing withA suitable private key was not found. (#860)- Docker images are published for every supported Node version (#760); wall-profiling workers that nothing captures are no longer started, and
@datadog/pprofis 5.18.1 (#821);npm run buildexits 0 and PRs are gated onnpm run typecheck(#818, #881, #901).
Already shipped in 5.2.x
These were backported to the 5.2 patch train, so a node already on 5.2.14 has them: connection state corrected on worker exit and node removal with link metrics and recovery-fire counts (#814); the shared status buffer anchored to survive its owning worker (#845); multi-hop dedup exclusion covering directional sendsTo peers (#809); a copy-progress watchdog that fires only on transport evidence (#782); a bounded source-503 blob-gap escalation budget (#797); computed fields covered by the initial copy (#772); and the one-origin-per-transaction apply fix in reduced form (#906, a port of #907).
Also in this release
Core submodule syncs to each 5.3 pre-release and to v5.3.0; release tooling fixes (projected core version in the dry run, stale draft-release Slack link) (#854, #855); cluster and replication test hardening and deflakes; review-coverage CI in enforce mode; release backports targeted by milestone; Docker base image, dependency and workflow-action updates.
Harper (core) v5.3.0
Harper 5.3.0 is the first stable release of the 5.3 line. These notes cover everything merged to main since the 5.2 line branched at v5.2.4, across 5.3.0-alpha.1 and beta.1–beta.4 — not only the changes since the last beta. They are written as the difference from 5.2.14: changes that were backported to the 5.2 patch train are listed once, under "Already shipped in 5.2.x".
Headline work: native full-text search, a native memory-mapped HNSW vector index, branched databases, cluster-wide record locks, a models.decide primitive, rollback to a previous deploy by deployment_id, and a long list of replication, transaction and storage correctness fixes.
Upgrade notes
Read these before upgrading a production node or cluster.
- RocksDB storage format is one-way. Dropped and recreated RocksDB tables now use generation-suffixed column-family names. A release older than 5.3 opens the bare names and sees those tables as empty, so do not downgrade a RocksDB node below 5.3 after upgrading. (#2604)
- npm installs that use S3 export/import must add
@aws-sdk/client-s3and@aws-sdk/lib-storage. The AWS SDK is now an optional peer dependency; without it those operations return 501 with the install command. The official Docker image still includes it. (#2609) - Restoring an
exclude_blobsbackup now requiresallow_engine_only. Such a restore previously succeeded with only a log warning and could leave records pointing at the wrong blobs. Scripted engine-only restores fail until the flag is added, including restores into a newtarget_database. (#2646) - Subscription catch-up delivers only the current version of each record by default. Pass
includeSuperseded: truetoTable.subscribeto receive retained history;rawEvents: truestill includes every event, and durable MQTT subscriptions still receive every publication. (#2767) searchByIndexand a custom index'ssearch()take an options object. A positional fifth argument tosearchByIndexnow throws aTypeErrorinstead of silently permitting a scan it meant to forbid. (#2187)- Startup waits at most
deployment.startupInstallTimeout(default 10 minutes) for component installs before opening listeners;0restores the unbounded wait. (#2784) - Requests Harper used to accept silently are now refused:
set_configurationwith unrecognized parameters (#2272); a deploy ordrop_componentwhose root-config changeHARPER_CONFIG/HARPER_SET_CONFIGwould undo (409, #2801);Table.clear()on a table with an active full-text declaration (501, #2616); and granting the legacycatchupoperation through a role allowlist (#2809). - Aliased databases with file-backed blobs now resolve blob paths to the first-assigned identity. A deployment whose blobs were stored under an alias that happened to load last may see them resolve differently; no migration ships with this change. (#2683)
- Every client certificate is verified once more after the upgrade (mTLS verification only), because the verdict cache key changed. (#2835)
Full-text search (new)
Tables can declare @fullText fields and query them with BM25-ranked search through the existing Table.search, REST and Operations API paths, with the usual authorization. (#2855; docs: documentation#691)
- Declared fields are projected into local Tantivy indexes (via the optional
@harperfast/fulltextpackage). RocksDB and the audit log stay the source of truth; the index files are rebuildable, are reused across restarts when compatible, and are replayed or rebuilt from the audit log before the index reports ready. (#2569, #2615, #2616) - Dropping a database or table fences and drains every worker's native handles, with durable markers so an interrupted drop is retried.
- Catalog-scale qualification continues in #2512.
Vector search: native HNSW
- A file-primary, memory-mapped HNSW backend built on a shared derived-index runtime (#2430, #2567). An eligible index uses it when its table has durable audit enabled (or declares
audit: true);nativePlane: falseopts out. (#2657) - Queries rerank only the rows they return and resolve hits from keys stored in the plane (#2665). Index lag is bounded and callers can wait for prior writes to be visible (#2658, #2676).
- About half the layer-0 memory: the new per-index
nativePlaneLayer0Capdefaults to 64 instead of a fixed 128, with recall@10 within about 0.5 points at ef ≥ 128.@harperfast/hnsw0.4.0 cuts visited-set memory 32x; the on-disk format is unchanged, so no reindex is needed. (#2701, #2722) - Filtered vector search answers indexed conditions from secondary indexes instead of decoding a record at every visited node. On a 20k-record test, 10%-selective filtered top-50 queries went from 7.0 ms to 4.3 ms with recall unchanged. (#2691)
Branched databases (new)
An application can declare branchedDatabases: [data] and get a private, durable fork of that database, reached through the same databases import, so several variants of an application can run against isolated copies of the same data. The fork lives at a deterministic path, so it is re-adopted on restart and resolves to the same place on every node. (#2352)
- Branches declare their own tables through
@table,ensureTableanddefineTable(#2523), clone the blob store by hard link (#2426), and are removed bydrop_component(#2517). - A branch no longer discards its latest writes when restarted after a crash (#2414). Branch identity is owned from disk, so a bad marker or stranded root is recovered instead of bricking the node, and
create_databasecannot take a branch's identity.
Record locks (new)
table.lock(id) gives an application an exclusive lock on one record, serialized across worker threads (#2462) and across the cluster, with amortized per-record ownership so a hot record does not pay a round trip per lock (#2498, #2667). Successor freshness is fenced across handoffs (#2613, #2627, #2663). An exhausted wait now reports what the lock's home actually answered, so 423 (held, retry) and 503 (coordination broken) mean what they say (#2685), and a conflict retry on a lock-protected write no longer overflows the stack (#2825).
Models: models.decide (new)
Beside embed and generate, models.decide takes program state and a small closed schema (enum, boolean, bounded integer, or one level of named fields) and returns the chosen value plus a probability distribution over the allowed values, so an application can act automatically above a confidence threshold and route the rest to a person. It scores from token log-likelihoods where the backend supports it, supports an @decide schema directive and durable decisions with recorded outcomes, and is served by a new decision backend kind under models.decision. (#2836, #2848) The models config block also hot-reloads (#2377).
Deploy and components
- Return to a previous release by
deployment_id. A deploy now keeps the release it replaces, sodeploy_component { project, deployment_id }goes back to any release still withindeployment.stagingRetention.maxCountwithout rebuilding or reinstalling. Activating the live id succeeds without a swap, so a partly failed activation converges by retrying the same id. (#2897) - Safer deploys: the replacement is built aside and validated before the swap (#2345); a build can be staged with
activate: falseand activated later (#2605); staged builds are bounded bydeployment.stagingRetention.maxCount(#2531); the swap retries through a transient rename lock (#2570); components are restored after a failed preparation (#2066). - Cluster deploys: components whose dependencies extend a built-in while loading (reflect-metadata, tsyringe, TypeORM, NestJS) deploy on peers under the default
freeze-after-loadlockdown (#2893); peers get a deadline so an unresponsive peer cannot hold the origin's deploy indefinitely (#2818). harper deployshows why a deploy was refused instead ofDeploy completed (no result payload)(#2817). One application's hung plugin load no longer times out every other application's (#2884). An invalid GraphQL schema fails on its parse error instead of gating every application for 30 seconds (#2432).- OIDC trusted publishing: deploy from CI with no stored credential (#2173). An application can run in its own dedicated worker thread (#2524).
set_configurationwrites the config file Harper actually boots from (#2796).
Security
- An operation is authorized against the principal that invoked it (#2217), and SQL honors the permission denial
processASTcomputes (#2202). - Revoked client certificates (mTLS verification, off by default): three paths that could accept a revoked certificate are closed — non-expiring cached verdicts on RocksDB, CRLs inside their grace period, and the first check after a CRL download on LMDB. TLS state is published transactionally, so a failed rebuild cannot downgrade below the last good configuration. (#2835, #2384)
- SSH private keys no longer reach the logs:
add_ssh_key/update_ssh_keypayloads were written verbatim to the operations and MCP audit logs. (#2741) - Role allowlists can grant component and configuration operations (
deploy_component,restart_service,set_configurationand 17 more) by their API names, and component-registered operations are grantable too. (#2809, #2260) - Scoped authentication tokens via an inline role on
create_authentication_tokens(#2176);liveSubscriptionAuthcan revoke one subscriber without ending its subscription (#2039); a rejected cookie or mTLS session becomes a negotiated 401 instead of a raw error (#2717); OIDC replay records expire on every node (#2823); refused operations log one line naming what was refused (#2816). - Users and roles are read through the record cache, ending cache-refresh storms when user or role changes replicate in. (#2776)
Replication and subscriptions
- Replicated transactions stay whole: a receiver keeps one open transaction per connection, so one peer's transaction can no longer be split by or absorb another's, and a failed transaction is held for replay instead of skipped. (#2800)
- An out-of-order write below the audit-retention floor no longer stalls replication (on a 16-node cluster it left two legs with zero commits for over 24 hours). (#2684)
- A record key beginning with
/no longer hangs the subscription broadcaster at 100% CPU (also present in 5.1 and 5.2). (#2689) A subscription that ends mid-delivery no longer stops delivery to the key's other subscribers, and a throwingend_txnlistener is contained. (#2913) - Cluster-origin table definitions are additive-only, and a table being created is invisible to catalog scans on other threads, so a peer can neither destroy local attributes nor announce a partial attribute list. (#2258, #2381)
- Non-bare-host node identities are rejected and IPv6 replication URLs are formed correctly (#2223);
sourcedFromcache-fill conflicts converge (#2065); replicated apply failures are reported before later events are consumed (#2630).
Transactions, storage and recovery
- A force-committed long transaction resolves only after its commit lands, and rejects if the commit failed. Previously a source-apply transaction held past the limit could report its writes committed while the commit was still in flight. (#2916)
- Multi-worker counter increments are no longer lost (#2259); same-key write order is preserved for explicit saves (#2561); commit conflict retries are bounded by the request's queue-time budget (#2459); a wedged commit reports the long-lived transaction holding it (#2473); transaction-handle leak and poisoning cases are fixed (#2086, #2128).
- Recovery: transaction-log replay fail-stops at a corrupt frame instead of applying a partial transaction (#2087); the restoring marker is published atomically, so a crash mid-restore cannot leave a half-purged database that loads as healthy (#2644); a restore fences every blob-root mutation while it runs (#2646).
- Integrity:
copy-dbno longer produces a silently corrupt copy (#2098); a backup captures only whole blobs (#2265); resolved attribute values are kept out of durable records (#2368); record version andtxnLogKeyare separate audit clocks (#2497, #2499). - RocksDB: dropping a table is a logical removal with deferred column-family reclamation, so interrupted creates and drops finish without blocking schema reconciliation (#2604); the
WriteBufferManagerno longer stalls every writer by default (#2492); a write-ahead retention floor is raised before every prune (#2458); rocksdb-js 2.10.0. - LMDB: closing two aliases of one environment no longer segfaults (#2766); audit cleanup no longer records a failed removal as removed (#2808).
- Harper boots when its volume is full (#2245) and refuses to boot when the data-version stamp was not recorded (#2398). Opt-in deflate compression for file-backed blobs (#2460).
Queries, HTTP and APIs
- The query planner uses storage-level range estimates (#2163, #2479);
sql.engine,allowFullScan,maxSortRowsandmaxHashRowsnow take effect (#2484); indexed array-element scans return each record once (#2493). - REST total-count pagination via
Prefer: count=andContent-Range(#2147); arrayPUTfollows collection semantics and is atomic per batch (#2097); newputoperation (#2347); the Operations API resolves@relationshipattributes (#2302). - Streaming responses fail visibly. Streaming REST responses end with a structured error frame instead of a truncated body (#2614), and on the uWS backend a failed stream aborts the connection instead of completing as a truncated
200(#2900). On Bun, a streamed response closes when the client asks (#2351). - Compressed JSON responses are faster: buffer responses use brotli quality 2 (as streams already did) instead of 11, taking a 242 KB response from 0.9–3.8 s to about 0.7 s on a 2-thread node. (#2899)
withNodeAdapterruns Node middleware (Next.js,compression,send) against a realWritable(#2528).- MCP:
tools/listadvertises only what the permission check allows,structuredContentis always a JSON object, and nocreate_*tool is invented for a Resource without a create verb. (#2753, #2754, #2405) - TypeScript:
Table.d.tsis emitted under TypeScript 6+, andTableResourceClass/TableResourceInstanceare exported from the package root. (#2904)
Operations and diagnostics
- Jobs from a Harper process that is gone are marked
ERRORat boot instead of reporting as running forever. (#2645) - Startup and workers: swallowed uWS listen failures are surfaced (#2112); HTTP starts correctly after a pre-ready worker restart (#2129, #2314); a low
threads.maxHeapMemoryno longer makes a node unbootable (#2290); worker respawns no longer block container restarts (#2316, #2363). - Diagnosis: a warning when a worker's event-loop utilization stays pinned at ≥ 0.99 for 30 s (#2882); a stuck worker's OS thread state is logged on ITC timeout (#2521); request-queue shedding (503) is logged (#2828); the startup banner reports the Harper version.
- Analytics aggregation resumes from the last raw record it rolled up, so
get_analyticsno longer permanently under-counts samples that arrived mid-cycle. (#2692) - Logging: the window before config load no longer drops every log line (#2467); external and component loggers inherit rotation config (#1880);
logging.rotation.maxSizeis enforced on write (#2475). - Smaller installs: the AWS SDK is optional and CLI prompts moved to lazily loaded
@inquirerpackages, about 36 MB less per npm install. (#2609, #2626) - Windows: intermittent
set_configuration500s (#2339), 8.3 short watch paths (#2309), deleted watch paths (#2364), and PID reuse during process-tree termination (#2890) are fixed.
Already shipped in 5.2.x
These 5.3 changes were backported to the 5.2 patch train (5.2.5–5.2.14), so a node already on 5.2.14 has them: mid-scope commit atomicity and self-committing-transaction scopes (#2239, #2291, #2325); continuous RocksDB audit retention (#2338); re-delivered deletes no longer echoing across a mesh (#2761); bounded boot replay (#2788, reduced port); resumable secondary-index backfill (#2539, #2543); fail-closed mTLS revocation checks (#2457); HNSW efConstruction auto-scaling (#2181); secret custody for boot-time installs (#2783); and get_status worker counting (#1952).
If you ran a 5.3 pre-release
- A durable (QoS 1+) MQTT session resumed on 5.3.0-beta.3 could skip records; fixed in beta.4 (#2874).
- A replicated deploy of a component named
harpercould deadlock peers from alpha.1 through beta.3 (#2818). - A cluster on beta.2 or beta.3 must upgrade every node before a deploy of a component using reflect-metadata converges (#2893).
Also in this release
Release packaging fails on TypeScript errors and holds the compiler off 6.0.x–7.0.x (#2886, #2770); promoted QA regression anchors and new integration coverage across blobs, audit history, MQTT, SSE, TTL, caching, restore and eviction; flake fixes across Windows, unit and integration suites; CI job/step timeouts, review-coverage and planning-receipt enforcement; DESIGN.md split into per-directory design notes; dependency updates (rocksdb-js 2.10.0, @harperfast/hnsw 0.4.0, alasql 4.19.1, mocha 12, non-major bumps).
Harper core v5.3.0 release · Core changelog v5.2.4...v5.3.0 · Full Harper Pro changelog v5.2.4...v5.3.0