Releases: JasperFx/marten
Release list
V9.17.0
High-water health check: opt-in autoRestart + heartbeat primary signal (#4986)
Builds on the detection-only check from 9.16 (#4984). Requires JasperFx 2.32.0 (#539), which this release rolls up to (#4987).
- Opt-in
autoRestart—AddMartenHighWaterHealthCheck(TimeSpan? staleThreshold = null, long minimumGap = 1, bool autoRestart = false). When the check is Unhealthy andautoRestartis on, it asks the local projection coordinator's daemon to restart the high-water agent's poll loop only — the mark is never advanced — capped to once per staleness window per database. The cycle is still reported Unhealthy so an alert still fires. Intended for Solo / leader nodes. - Heartbeat is now the primary staleness signal — when
EnableExtendedProgressionTrackingis on, the high-water agent stamps a liveness heartbeat on theHighWaterMarkrow every poll cycle. Heartbeat age proves the loop is cycling independent of whether the mark advances, so a quiet store is never a false positive, and a dead agent is caught even when projections are fully caught up (the exact #4961 blind spot). The original sequence-gap heuristic is retained as theExtendedProgression-off fallback.
Full changelog: V9.16.1...V9.17.0
V9.16.1
Async daemon data-safety release: the high water detection can no longer advance past "outstanding" event sequence numbers — sequences reserved by transactions that are still in flight — which could silently skip those events in async projections under concurrent append load (bulk imports being the classic case). Root-caused and fixed from discussion #4953.
The four closed mechanisms:
- The
GapDetectorcommand batched three statements, each reading its own READ COMMITTED snapshot — commits landing mid-command could defeat every gap check and silently advance the mark over an in-flight append, regardless ofStaleSequenceThreshold. Detection is now a single statement / single snapshot. - Projection rebuilds and forced catch-up looped the gap-skipping detection toward the reserved sequence
last_value, mowing through in-flight gaps.CheckNowAsync(JasperFx.Events 2.29.1) now targets the highest committed sequence and simply waits for in-flight appends to land. - The stale fallback could teleport the mark to
reserved last_value - 32across thousands of in-flight reservations on an idle-then-suddenly-busy store, because its gate measured staleness againstmt_event_progression.last_updated. The threshold is now measured from when each specific gap was first observed. - Wall-clock stale skipping could not tell a slow transaction from a rolled-back one. Before skipping any stale gap, Marten now checks PostgreSQL for evidence that a transaction which could still fill the gap is alive (
pg_lockson the mt_events tables, open transactions inpg_stat_activity, in-progress write xids frompg_current_snapshot()), and holds while any exists — by default Marten never knowingly skips past a live appender. Only provably-dead gaps (rolled-back appends) are skipped, bounded to the sequence ceiling observed with the gap, and every skip is logged at Warning with its exact range.
New knobs on StoreOptions.Projections: UseTransactionEvidenceForGapSkipping (default true; false restores the previous wall-clock behavior) and SkipStaleGapsDespiteLiveTransactionsAfter (default null = never skip a live appender; PostgreSQL's idle_in_transaction_session_timeout is the recommended backstop against leaked sessions).
What's Changed
- fix(#4953): high water detection never crosses outstanding event sequences by @jeremydmiller in #4977
- Consumes JasperFx.Events 2.29.1 (JasperFx/jasperfx#530)
Full Changelog: V9.16.0...V9.16.1
V9.16.0
Lot of CritterWatch, couple bug fixes too
What's Changed
- Fix AdvancedSql/raw-SQL scalar queries for reference-typed columns (byte[], IPAddress, etc.) by @mdissel in #4960
- fix(#4961): PostgresqlListenWakeup falls back to a timeout wait when the DB is unreachable by @jeremydmiller in #4965
- Bump JasperFx to 2.28.0; declare EventProjection doc types for rebuild teardown (#4685 COPY) by @jeremydmiller in #4969
- fix(#4966): update natural key on projection rebuild (JasperFx 2.28.1) by @jeremydmiller in #4970
- test(#4963): verify + document the blue/green side-effect gate by @jeremydmiller in #4971
- fix(#4964): hold the Normal high-water mark before a leading sequence gap by @jeremydmiller in #4972
- refactor(#4968): route stream archive through the shared Weasel event-store seam by @jeremydmiller in #4973
- feat(#4962): targeted per-cell ReadProjectionProgressAsync on MartenDatabase by @jeremydmiller in #4974
- feat(#4975): exact ReadProjectionProgressAsync(ShardName) override + JasperFx 2.29.0 by @jeremydmiller in #4976
Full Changelog: v9.15.4...V9.16.0
Marten 9.15.4
What's Changed
- Support Where Any on collections when using strongly typed Ids by @ximon in #4957
- Bump ws from 8.19.0 to 8.21.1 by @dependabot[bot] in #4955
- #4956 — lock down ForTenant() conjoined-document tenant isolation by @jeremydmiller in #4958
New Contributors
Full Changelog: v9.15.3...v9.15.4
Marten 9.15.3
This addresses a potential vulnerability from SQL injection via non-string constant in a LINQ Select projection
Not a common usage, but still.
What's Changed
- Parameterize non-string Select projection constants (SQL injection fix) by @jeremydmiller in #4954
Full Changelog: 9.15.2...v9.15.3
Marten 9.15.2
Marten 9.15.2
A patch release. Both fixes come out of the same 512-tenant-database production deployment, reported by @erdtsieck, and both turned out to be worse than the reports described.
Bulk event insert ran a full schema apply on every batch
The batch BulkInsertEventsAsync overloads opened with Storage.ApplyAllConfiguredChangesToDatabaseAsync() on every call.
That is not a cheap check. It calls Tenancy.BuildDatabases() and runs a full schema delta — partition introspection plus information_schema sweeps — across every database in the store. So a sharded store paid one apply per database, per batch. On the reporting deployment, each ~1,000-event batch was triggering 512 schema applies.
The measured effect: import throughput collapsed to ~17 events/s, against >3,000/s for the streaming overload. A 686k-event tenant projected to roughly 11 hours. The connection pool filled with ~370 backends whose last statement was Weasel's partition-introspection query, which fed directly into the server-wide connection pressure that deployment was already fighting.
That the streaming overload BulkInsertEventStreamAsync has no such call and is fine is the tell: the schema apply was never part of the contract. It was a leftover.
The apply is now:
- skipped entirely when the effective
AutoCreateisNone— it is a no-op there by contract, so all that remained was the introspection cost; and - otherwise run at most once per database the import actually touches, memoized on
IMartenDatabase.Identifier.
One subtlety worth recording, because it is the kind of thing that bites later: the memoized apply deliberately does not take a caller's CancellationToken. The first caller to arrive owns the single in-flight task that every concurrent caller for that database awaits — so binding that shared task to one caller's token would let a single cancelled batch fail sibling batches that were never cancelled. Each caller applies its own token at the await site instead. A schema apply is short and idempotent, so letting it run to completion is the cheaper trade.
Under AutoCreate.None, the event storage must already exist before import. That is the documented contract and it matches the streaming overload — but if you were previously relying on the per-call apply to create it for you under a non-None store, note the change.
The document bulk-insert path (BulkInsertAsync / BulkInsertDocumentsAsync) is unaffected. It routes through the ordinary per-feature EnsureStorageExistsAsync that Weasel already memoizes, not a full-store delta.
Tenant provisioning silently under-provisioned partitions
AddPartitionToAllTables, and the tenant-provisioning paths built on it, walked the calling store's StoreOptions to decide which tables needed a list partition for a new tenant.
So any tool or host that provisions tenants from a store which doesn't register every document type silently under-provisioned. Document types unknown to the caller never got their partitions — and the tenant then failed with a Postgres 23514 check-constraint violation on first write to the missing partition. Nothing failed at provisioning time; the damage surfaced later, somewhere else.
The workaround was "the provisioning tool must register all document types," which re-creates schema knowledge in a second place and drifts as document types are added.
The sweep is now database-driven: it enumerates tenant list-partitioned tables from the Postgres catalog, so a partially-registered store still provisions every partitioned table it finds.
Scoping is enforced inside the catalog query rather than filtered in memory afterward:
- Schema — the store's own
AllSchemaNames()only. Foreign partitioned tables in a shared database are never touched. - Partition shape — LIST strategy, exactly one key column, and that column named
tenant_id. This is the filter that matters most, and it is what keeps the sweep off Marten's own non-tenant list partitioning:UseArchivedStreamPartitioningkeysmt_eventsonis_archived, andByList()keys on its own field. Without it, a "helpful" sweep would start adding tenant partitions to tables partitioned on something else entirely. - External management — tables marked
ByExternallyManagedListPartitions()are subtracted.
Opt out with SweepPartitionedTablesFromDatabase (default on). No Weasel change was required.
Known limitation, and it is a real one: a document type registered into a schema the calling store has never heard of stays invisible to the schema filter — a store cannot own a schema it does not know exists. Single-schema stores (the default, and the reporting deployment's shape) are fully covered. Closing this properly would need a persisted table list alongside mt_tenant_partitions.
Everything in this release
| Issue | Pull request | Summary |
|---|---|---|
| #4946 | #4949 | Bulk event insert applies the schema once per database, not per call |
| #4944 | #4950 | Database-driven tenant partition sweep |
| — | #4951 | Release prep: 9.15.2 |
Full changelog: 9.15.1...9.15.2
Marten 9.15.1
Patch release for a silent data-correctness regression. If you use ForTenant() on an identity-mapped or dirty-tracked session, upgrade.
Fixed
-
#4947 —
ForTenant()on an identity session stopped returning tenancy-neutral documents (reported by @dervagabund, with a repro — thank you). AForTenant()view of an identity- or dirty-tracked session no longer saw global (tenancy-neutral) documents tracked by the parent session. Since a global document has exactly one row per id for the whole database,LoadAsyncthrough theForTenantview missed the identity map, went to the database, and returnednullfor a document that is there. A silent wrong answer, not an error.Affected: 9.13.0, 9.14.x, 9.15.0. Introduced by the fix for #4801, which tenant-scoped the identity map and version tracker for
ForTenantsessions. That was correct for conjoined documents — where the same id means a different document per tenant — but it was applied per session rather than per document type, so it also isolated document types that are tenancy-neutral and must be shared.Sharing is now decided per document type. A nested
ForTenantsession shares the parent's identity-map and version-tracker entry for a type only when the storage is identity-mapped, the type is notConjoined, and the nested session's database is the same instance as the parent's (under database-per-tenant, the same id in another tenant's database is a different document even for a tenancy-neutral type). The isolation introduced by #4801 is preserved exactly — theBug_4801suite still passes, and the new tests include guard rails asserting conjoined documents stay isolated.
Full changelog: 9.15.0...9.15.1
Marten 9.15.0
Closed issues
- #4942 — sharded tenancy: auto-assign never repaired half-provisioned tenants (PR #4945).
findOrAssignTenantDatabaseAsyncreturned early on an existing assignment row, skippingcreatePartitionsForTenant+ per-tenant event-sequence provisioning — so a tenant whose provisioning was interrupted (assignment committed, partitions missing) failed every write with23514forever. Both early-return paths (including a second race-window hole under the advisory lock) now run the same idempotent repair the explicitAddTenantToShardAsync(tenantId, databaseId)overload always ran, guarded to once per process per tenant via the resolution cache. - #4941 — two-day silent projection outage (closed with full mapping). Root cause was #4942; the invisibility was JasperFx/jasperfx#506/#507, fixed in JasperFx 2.27.0 which this release consumes.
Also in this release
- Bundles the fixed
JasperFx.Events.SourceGeneratoranalyzer (JasperFx/jasperfx#505) — CS1061 compile break for no-parameterless-ctor aggregates with instanceApplyreturning the aggregate. - Follow-up enhancement filed as #4944 (database-driven partition sweep via
pg_inherits) for the #4943 provisioning-tool scenario.
Verified against Wolverine (full solution + CoreTests/MartenTests/distribution/Http suites, zero failures) and CritterWatch before publishing. Thanks to @erdtsieck for the dump-verified root-cause analysis.
Marten 9.14.1
Marten 9.14.1 is a patch release focused on a substantial round of LINQ query-translation improvements, plus event-store partitioning, high-water, and AoT fixes, and refreshed Weasel/JasperFx dependencies.
LINQ query translation
This release significantly expands what the LINQ provider can push down to PostgreSQL instead of falling back to slower strategies or throwing:
- Collection
Any(predicate)filters now translate to JSONPath and OR-of-containment strategies, and the old explode/ctidfallback has been replaced by a correlatedEXISTSstrategy.All()shapes and duplicated array fields moved onto the sameEXISTSstrategy. The net effect is correct, index-friendlier SQL for nested-collection predicates. - Indexing into complex child collections inside
Where()clauses is now supported (e.g.x.Children[0].Name == "..."). - Aggregates over collections —
Sum/Min/Max/Average— can now be used insideWhere()clauses. Regex.IsMatch()is translated inWhere()clauses.IComparable.CompareTo()now works for non-string comparables such asGuid(#4920), alongside broaderCompareTo()coverage,stringIsOneOfvia the?|operator, andCollectionIsEmptyviaICollectionAware.GinIndexJsonDataMember()was added for member-scoped expression GIN indexes.
#4916 — subclass queries now use duplicated fields and the base id
Querying a document subclass and filtering on a Duplicate()'d field or the base-class id previously emitted a JSONB filter (CAST(d.data ->> 'FarmId' as uuid)) instead of the real column, missing the duplicated column and the primary-key index:
o.Schema.For<Animal>().AddSubClass<Cow>().Duplicate(x => x.FarmId);
Query<Cow>().Where(x => x.FarmId == id) // now: d.farm_id = :p0 (was: CAST(d.data ->> 'FarmId' ...))
Query<Cow>().Where(x => x.Id == id) // now: d.id = :p0 (was: CAST(d.data ->> 'Id' ...))A subclass shares its parent's table, so the parent's column-backed members (duplicated fields, the id, the soft-delete flag) are now inherited by the subclass's query member resolution. Querying the parent type was already correct and is unchanged.
Event store, partitioning & daemon
- #4924 — hyphenated / GUID tenant ids under
UseTenantPartitionedEvents. Registering a tenant whose partition suffix contains a-(so every GUID tenant id) madeApplyAllConfiguredChangesToDatabaseAsync()throw42601because the per-tenantCREATE SEQUENCE/DROP SEQUENCEDDL emitted the identifier unquoted. The schema-apply statements are now quoted (matching the quick-append function and the imperative provisioning path), so hyphenated tenants migrate cleanly. Quote — not sanitize — so the append function can still resolve the sequence by its raw suffix. - #4915 — projection coordinator shutdown. The projection coordinator now drains on disposal, and via the Weasel 9.16.3 bump the advisory-lock
ObjectDisposedExceptionpath latches-and-rethrows so a HotCold cold node's leadership loop terminates instead of re-polling a disposed data source during shutdown. - #4913 — high-water scan under partitioning (JasperFx 2.26.0). Under
UseTenantPartitionedEventsthe store-global high-water agent was continuously runningselect max(seq_id) from mt_events, an unfiltered scan that fans out across every tenant partition on every poll. That store-global mark is not used to advance tenant projections (they advance per-tenant), so the recurring scan is now skipped under partitioning; tenant high water is driven by the per-tenant coordinator and poll timer. - #502 (#4922) —
GetProjectionStatusesAsyncnow resolves the correct named database.
AoT / trimming
- #4917 — corrected AoT annotations in the event graph.
- The
AddEventType/QueryRawEventDataOnlygeneric-constraint tightening was reversed, and event-mapping construction now routes through the cachedGenericFactoryCachewhile preserving the trimming root (#4930).
Dependencies
- Weasel 9.16.3 (#4932) — advisory-lock disposed-pool fix (marten#4915).
- JasperFx 2.26.0 — the #4913 high-water fix, plus 2.25.0's
ShardState.DatabaseIdentifier(#501).
Closed issues
Marten 9.14.0
Marten 9.14.0 is the recommended upgrade for all 9.x users. It combines the LINQ SQL-injection security fix (first shipped in 9.13.0) with the fix for the projection-coordinator shutdown race in #4874 and the accompanying dependency updates.
Beyond the LINQ updates, this made the new Per-Tenant Event Partitioning much more robust as we're testing that in conjunction with a JasperFx client for ludicrous scalability.
🔒 Security — SQL injection in the LINQ provider (GHSA-rfx3-98h7-v3xp)
Several LINQ / tenant-management code paths interpolated a runtime, potentially attacker-influenced value into generated SQL as a single-quoted literal without escaping or parameterization. A value containing a single quote could break out of the literal and inject SQL. The primary vector — a Dictionary<,> indexer key in a Where filter (a common "filter by attribute name" / EAV pattern) — was reported privately with an executed proof-of-concept and enabled filter / multi-tenant authorization bypass and blind data exfiltration.
Fixed sinks (#4911):
DictionaryItemMember— dictionary indexer key, e.g.Where(x => x.Attributes[key] == v)DictionaryContainsKeyFilter—Dictionary.ContainsKey(key)(Newtonsoft serializer + the Enum branch, which bypass System.Text.Json's quote escaping)SelectParser— a constant string projected throughSelect(x => new { L = runtimeString })DeleteAllForTenant— tenant id reaching per-tenant projection teardown (now parameterized)DatabaseScopedTenantPartitions— tenant id inlined into partition DDLEventLoader— per-tenant partition-pruning literal (defense-in-depth)
Each sink now escapes embedded single quotes or binds the value as a parameter; regression tests lock down every vector, and a follow-up LINQ-wide audit cleared the rest of the query hot path (full-text search, string-method translations, comparisons, IsOneOf/Contains/subset operators, and patching paths). Affected versions: 7.0.0 – 9.12.0. Also patched in 8.37.4 (8.x line) and 9.13.0.
Reported responsibly by @svenclaesson — thank you. See advisory GHSA-rfx3-98h7-v3xp (CVE pending assignment).
🛠️ Reliability — projection-coordinator shutdown drain race (#4874)
On host shutdown, the native HotCold projection coordinator could abort with ObjectDisposedException: 'Npgsql.PoolingDataSource' — the coordinator's leadership poll issued an OpenAsync against an already-disposed data source while tenancy was tearing down. This is the "case B" ordering storm reported against #4874 (distinct from the async-tenancy foundation laid in #4907, which did not resolve it).
The fix ships through the dependency updates below, with a Marten-side regression test (Bug_4874_coordinator_drain_ordering, #4912):
- JasperFx 2.24.1 (#499/#500) —
ProjectionCoordinatorBaseterminates the leadership loop on a disposed data source / wrapped cancellation instead of re-polling. - Weasel 9.16.2 (weasel#349/#350) —
AdvisoryLockguards against a disposedNpgsqlDataSourceduring shutdown (short-circuits while disposing and treats a disposed-poolObjectDisposedExceptionas a non-acquire rather than propagating).
⬆️ Dependency updates
- JasperFx 2.24.0 → 2.24.1
- Weasel 9.16.1 → 9.16.2
- Weasel.EntityFrameworkCore 9.2.1 → 9.16.2 (released from its prior version hold now that the Weasel line is published)