Releases: Hikari-Systems/simmer
Release list
v0.9.0
v0.9.0
Breaking. Two features: named ramps (D-099) and domain groups by MX host
(D-100). Config, admin paths, metric labels and the database schema all change.
Read the upgrade notes before rolling this out.
One process can now warm several unrelated programmes, each as its own ramp: a
complete routing profile with its own domain groups, sender rules, routes,
schedules and quota. Ramps share nothing, and route names are scoped to their
ramp.
A message's ramp is chosen at the final dot (§5.8), in this order:
- The listener's
ramp, unless the listener setsheader_overrides_affinity. - An
X-Simmer-Rampheader naming a ramp in the user'sgrants.ramps. - The listener's
ramp. default_ramp.
An unusable header is ignored and counted, never refused. The header is always
stripped before relay.
A group may list mx host suffixes. A recipient domain that no group lists
literally joins the group whose suffix matches its lowest-preference MX host.
So a Google Workspace company domain now counts against google, and holding
google at 0 holds it too. A failed or slow lookup (over 2 s) falls back to the
catch-all and never defers mail.
- Config: move
domain_groups,senders,default_chain,
strict_senders,thread_affinity,exhausted_chain_replyandroutes
underramps.<name>, and adddefault_ramp: <name>. An old config is refused
at startup, naming each key that moved. - Stop-the-world: stop every v0.8 instance, start one v0.9 instance, then
start the rest. Take a backup first; there is no downgrade.- At first start, existing quota, pause and frequency state is adopted into
default_rampbefore any listener binds. - A v0.8 binary cannot use the migrated schema: on Postgres it refuses to
start; on SQL Server every reservation fails with 451. Nothing miscounts.
- At first start, existing quota, pause and frequency state is adopted into
- Admin API: route paths move under
/ramps/{ramp}/…, and the old
/routespaths answer 410./quota/resetnow requiresramp. - Metrics: every route-labelled series gains a leading
ramplabel. Add
rampto anysum by (route). - SQL Server: route, domain group and ramp names are now limited to 128
characters. A longer stored name fails the migration.
Images, built from 35fffb148f80c39df05fdb8150582edebcc9e310:
- Postgres:
hikarisystems/simmer:v0.9.0(linux/amd64, linux/arm64) - SQL Server:
hikarisystems/simmer:v0.9.0-mssql(linux/amd64, linux/arm64)
Full Changelog: v0.8.0...v0.9.0
v0.8.0
v0.8.0
A feature release: the partial ramp can now compute its own share. Off
unless configured — a configuration that does not set share, or that sets the
D-091 list, behaves exactly as on v0.7.2.
warmup:
started: "2026-08-01T09:00:00Z"
schedule:
default: [50, 100, 200, 400, 800]
share:
mode: auto
floor: 0.05 # never offer less, until the tail releases it
ceiling: 1.0 # never offer more — the only promise about the peak
gain: 4.0 # how sharply it reacts to running ahead of pace
fill_by: 0.6 # meet the cap 60% of the way through the ramp day
tail: { below: 0.1, ceiling: 1.0 } # release the last tenthEvery parameter is optional and the values above are the defaults, so
share: {mode: auto} on its own is the whole feature.
Why. The share list added in v0.6.0 is a hand-tuned guess at one number
— the cap divided by the volume the chain carries — and it is wrong whenever
that volume changes. Set it too low and the cap never fills, so the ramp
quietly falls behind the schedule it is supposed to be following. Set it too
high and the day is a burst again. Nobody knows the right number on day 0, and
it moves.
How it works. The share is computed per (route, domain_group) from the
day's quota row and the clock: clamp(floor, ceiling, (cap_left / window_left) ^ gain). Run ahead of pace and it collapses; fall behind, or run
out of window, and it opens back up to the ceiling. Which messages are offered
is still the keyed hash, so every instance and POST /dryrun agree; the cap
itself is untouched; a pinned thread-affinity reply is still exempt; and a
graduated route is still offered everything.
Two things to understand before setting the knobs.
fill_by is below 1 on purpose. Simmer never delays a message — one it does
not offer to the warming route goes down the chain immediately — so it cannot
smooth a bursty sender, only take a fraction of each burst. Pacing to land
exactly at the day boundary loses the cap whenever the traffic is front-loaded:
the ramp throttles the morning to save room for an afternoon that never comes.
Aiming at an earlier deadline means hitting the cap early rather than not at
all.
The last of the cap is released, not floored. A floor does not finish a
ramp: at 0.05, the last 40 messages of an 800 cap need some 800 offered
messages to arrive. Under tail.below of the cap the throttle comes off
entirely.
What you may notice. A single early burst does not fill the cap, and is not
meant to — an hour into the ramp day, with the whole day ahead, taking the cap
at once is the burst this feature exists to prevent. A day whose traffic all
arrives in the morning will end short of its cap; that mail is not lost, it
goes out on the next route in the chain. ceiling is the one hard promise
about the peak a provider sees, and it defaults to 1, which promises nothing.
Observability. /routes reports mode, the parameters, and the share in
force per domain group (the cap it paces against is per group, so the
route-level today is null under auto). simmer_partial_ramp_share{route, domain_group} is a new gauge, absent when every message is offered. Watch it
against simmer_route_skipped_total{reason="partial_ramp"}.
Cost. One extra non-locking SELECT per message per auto route the walk
evaluates. A listed share, and every route with no partial ramp, read nothing.
Also in this release: the -mssql image went multi-arch in v0.7.2.
See DECISIONS.md D-097, docs/SOAK.md §15, and README "A share that tunes
itself".
Images, built from 97b701c11240a7bf34c524bc5c716e637b0c143f:
- Postgres:
hikarisystems/simmer:v0.8.0(linux/amd64, linux/arm64) - SQL Server:
hikarisystems/simmer:v0.8.0-mssql(linux/amd64, linux/arm64)
Full Changelog: v0.7.2...v0.8.0
v0.7.2
v0.7.2
The -mssql image is now multi-arch. hikarisystems/simmer:v0.7.2-mssql
carries both linux/amd64 and linux/arm64, so D-084's SQL Server build runs
on Apple silicon and on arm64 servers without emulation. The default Postgres
image was already multi-arch and is unchanged.
Nothing else changes. No configuration, no behaviour, no schema — this release
is the same code as v0.7.1 for every existing deployment.
tiberius was never the blocker. It is a pure-Rust TDS client and compiles
for arm64 unchanged; its one native dependency is OpenSSL, which the Dockerfile
already carried on both stages. The actual gate was the test side: Microsoft
publishes no arm64 SQL Server. Every 2025 tag of mssql/server is a
single-platform amd64 manifest, verified against the registry rather than
assumed.
The arm64 leg therefore runs its tests against Azure SQL Edge — the same
engine, published multi-arch — and was verified at 1054 passed, 0 failed
before being adopted. Note that Azure SQL Edge was retired on 2025-09-30, so
that leg is running on an image with no future updates; if it becomes
unusable, the arm64 test leg goes before the arm64 image does.
Babelfish was rejected as the alternative: it accepts and silently ignores
table hints, so UPDLOCK, SERIALIZABLE — the lock that makes §7.4's
reservation protocol correct — would have passed its tests while not
serialising anything.
See DECISIONS.md D-096.
Images, built from 15f0f2a4ac9cd1aff59ca9ef66a070f211be4d1a:
- Postgres:
hikarisystems/simmer:v0.7.2(linux/amd64, linux/arm64) - SQL Server:
hikarisystems/simmer:v0.7.2-mssql(linux/amd64, linux/arm64)
Full Changelog: v0.7.1...v0.7.2
v0.7.1
v0.7.1
A security fix. Upgrade needs no configuration change.
Simmer ended a DATA transfer on <LF>.<LF> as well as on <CRLF>.<CRLF>.
RFC 5321 §4.1.1.4 says the end-of-data marker is CRLF . CRLF and nothing
else, and the gap between those two readings is SMTP smuggling: a sender
that honours only CRLF.CRLF writes <LF>.<LF> inside one message, Simmer
ended the message there, and everything after it was read as fresh commands —
a second MAIL FROM / RCPT TO / DATA, with recipients the client's
envelope never named, relayed under the route's warming identity.
That identity is the asset the product exists to build, which is what makes
this shape worse here than in a general-purpose MTA.
What you may notice after upgrading. A . line now ends DATA only when
its own line ending and the previous line's are both CRLF. Any other spelling
is consumed as data and the whole message is refused 554 5.6.0 — never split,
never partly relayed. If one of your applications emits bare line endings, its
messages stop going through where they previously did. simmer_ambiguous_terminator_total
counts every occurrence; alert on it. A steady low rate is usually a broken
client of your own, and anything else is an attempt to inject a second envelope.
Bare LFs inside a message body are unaffected: they are still accepted and
still normalised to CRLF, exactly as before.
Also refused: a line that would reach the downstream carrying <CR>.<CR>.
Simmer has no bare-CR line semantics, but it forwards bare CRs verbatim, and a
receiver that honours them would split the message on the way out instead. The
check runs over the whole line rather than its content, because strip_eol
takes the line's own CR and the CRLF restored on transmit puts it back — the
first cut of this fix scanned the content and missed exactly the spelling the
published attack uses.
SPEC.md is amended: §5.5 (the end-of-data rule), §9.1 (the counter), §10.3
(the one permanent reply that is a statement about the message, and why §14.1's
worked example does not reach it). The reasoning is DECISIONS.md D-095, and
tests/finding_smtp_smuggling.rs is the regression suite.
Co-Authored-By: Claude Opus 5 (1M context) noreply@anthropic.com
Images, built from c2eb9d6299df14047424546612c5ba99fd955812:
- Postgres:
hikarisystems/simmer:v0.7.1(linux/amd64, linux/arm64) - SQL Server:
hikarisystems/simmer:v0.7.1-mssql(linux/amd64)
Full Changelog: v0.7.0...v0.7.1
v0.7.0
v0.7.0
Breaking for monitoring: /metrics is now off unless you turn it on. To keep
it, add one line to your config before upgrading:
admin:
listen: "0.0.0.0:8080"
auth_token: "${SIMMER_ADMIN_TOKEN}"
metrics: true # or { enabled: true, idle_timeout: 24h }A config that doesn't mention admin.metrics still loads, but /metrics
answers 404, so Prometheus scrapes and every alert built on simmer's metrics
stop. Startup logs a warning saying so on every boot. metrics: false turns it
off deliberately and silences the warning. Mail handling, routing, quotas and
the rest of the admin API are unchanged.
Why. Every series the exporter holds exists only to answer a scrape, and one
of them grew without bound: simmer_unmatched_sender_total{domain} added a series
for every distinct sender domain that matched no rule, and kept it for the life of
the process. Measured on the soak at about 204 bytes per domain (r = 0.979),
it was simmer's only real memory growth. A daemon fed new sender domains would
reach about 195 MiB per million, and nothing recovered it short of a restart.
What changed:
- Off (the default): no metrics recorder is installed at all, so nothing is
held and every metrics call is a no-op. - On: a counter nothing has incremented for
idle_timeout(default 24h, at
least 1m) is dropped, and comes back from zero on its next increment, which
Prometheus'srate()andincrease()read as an ordinary counter reset. That
bounds the unmatched-sender growth by time. It applies to every counter, so a
route silent for a day loses itssimmer_messages_totalseries until its next
message. The exporter prunes while rendering, so enable it only where
something scrapes it.
SPEC.md is amended (§4.1, §4.2, §9.1, §9.2); the reasoning is DECISIONS.md
D-093.
Also in this release:
- No
unsafein simmer's own code, enforced (D-094):
[lints.rust] unsafe_code = "forbid"over the library, the binaries and every
test. There was none to remove; now there can't be. Dependencies are outside
any crate's lints and still use it (tokio, rustls, libc, jemalloc). - Allocator counters for the soak tier (D-092): an
alloc-statscargo
feature writes jemalloc's own counters to a file the soak samples. No
published image enables it. It showed that the soak's memory gate has been
judging jemalloc's page cache rather than simmer (live heap 1.5–2.0 MiB while
cgroupanonswings 5–70 MiB):docs/SOAK.md§13. - CI: the image builds cache their Docker layers in the GitHub Actions cache.
Verified: the full suite on both builds (Postgres 1198 passed, SQL Server 1046
passed), clippy on the default, alloc-stats and mssql builds, cargo-deny on the
two published builds, an hour's soak with the allocator counters on (simmer's live
heap +0.39 MiB/h, all of it the unmatched-sender series), and 30 minutes with
metrics off, where it was flat (−0.05 and −0.01 MiB/h) and the soak's F7 check
passed for the first time. Not run: the acceptance tier.
Images, built from ea8837d937e794bad1927aca7e39ce07bf1485fb:
- Postgres:
hikarisystems/simmer:v0.7.0(linux/amd64, linux/arm64) - SQL Server:
hikarisystems/simmer:v0.7.0-mssql(linux/amd64)
Full Changelog: v0.6.0...v0.7.0
v0.6.0
v0.6.0
A feature release: the partial ramp. A warming route can now be offered only
a share of its traffic for its first days, so its cap fills later in the day
rather than in a burst after each day boundary. Off unless configured: a
configuration that does not set it behaves exactly as on v0.5.0.
warmup:
started: "2026-08-01T09:00:00Z"
schedule:
default: [50, 100, 200, 400, 800, 1500]
share: [0.1, 0.25, 0.5] # day 0: a tenth of the messages try it; day 3 on: allWhy. A warming route took every eligible message from the day boundary
until its cap was met, and then none. Early in a ramp, when caps are smallest,
the whole day's volume went out in the first hours: a burst, not a steady rate.
How it works. share is indexed by day index like the caps, and applies to
every domain group. On day i, only share[i] of the messages that reach the
route in the walk are offered to it; the rest steer to the next link, exactly as
recipient_frequency does, and nothing is dropped. The cap itself is unchanged.
Past the end of the list every message is offered. Unlike a cap, the last value
does not repeat, because a list ending below 1 would otherwise throttle the
route forever.
Which messages are offered is a keyed hash (the recipient-frequency salt over the
normalised recipient, the route and the day index), not a dice roll. Every
instance agrees without shared state, and POST /dryrun reports the real answer
for a recipient. A pinned thread-affinity reply and a graduated route are exempt.
Refused at startup (§4.2): a share outside (0, 1], and a route with a
share list that is last in any chain, where every message it turned away would
be a 451.
Observability. simmer_route_skipped_total{reason="partial_ramp"}, a
name=partial_ramp step in the chain-evaluation log line and in dry run, and
partial_ramp: { share, today } in GET /routes.
SPEC.md is amended (§3.2 step 3c′, §4.1, §4.2, §7.2, §9.1, §9.2, §9.4); the
reasoning is DECISIONS.md D-091.
Verified: the full suite on both builds (Postgres 1188 passed, SQL Server
1037 passed), clippy and cargo-deny on both, and a live run on the acceptance
stack. Not run: the acceptance tier, which has no partial-ramp case.
Images, built from 1fd176bea4ddfb253b6816a2077219fd0ab997e3:
- Postgres:
hikarisystems/simmer:v0.6.0(linux/amd64, linux/arm64) - SQL Server:
hikarisystems/simmer:v0.6.0-mssql(linux/amd64)
Full Changelog: v0.5.0...v0.6.0
v0.5.0
v0.5.0
A feature release: thread affinity. A reply the application sends into a
conversation Simmer started now leaves via the route that started it — the
identity the recipient has already seen — instead of whichever route today's ramp
would pick. Off by default: a configuration that does not set it behaves exactly
as on v0.4.0.
thread_affinity: true # top level, beside strict_sendersWhy. Every message was routed afresh. A recipient who replied to
sales@newbrand.com in the morning heard back from news@mail.established.com
in the afternoon, once the warming route had spent its day — or once their own
back-and-forth had tripped its recipient_frequency threshold.
How a reply is recognised. Nothing is stored. Each route's Message-ID:
carries its own domain, so an ID Simmer emitted names the route that emitted it.
The recipient's mail client puts that ID in In-Reply-To:/References:; the
application carries References: forward when it answers, and Simmer reads it.
The README's "Thread affinity" section lists the headers the application must
build.
- Past the cap — only once it has been met, and counted. A pinned reply takes
an ordinary slot first, so replies within the cap spend it like any message.
When the cap is already met, the reply still goes out on its route, past the
cap:committedreads aboveallowance, the allowance itself never changes,
and every message that is not a pinned reply is still refused at it. - Past
recipient_frequencyon the pinned route. The event is still
recorded. - Not past a pause, a strict preflight failure, or a
warmup.startedin the
future: the reply then takes the ordinary walk. A downstream failure is still a
failure — no failover. - Refused at startup with it on: a route in any chain whose
Message-ID:
has no literal domain, and two routes in one chain sharing one. - Envelope-only configurations no longer reject at
RCPT TOwhile it is on:
whether a message is a reply is in its headers. - New metric:
simmer_thread_affinity_total{route,outcome}—hit,
over_cap,ineligible,unmatched. Watchover_cap: each is a send the
ramp did not schedule.POST /dryruntakesin_reply_toandreferences
and shows the pin.
Now in the specification. SPEC.md is amended: §2.1, §2.2 (clarified), §3.2
step 2a, §4.1, §4.2, §5.4, §7.3, §7.4 — its one exception to "overshoot is not
acceptable" — and §9.1, §9.4, §9.5. DECISIONS.md D-090 records the reasoning and
the alternatives rejected; open question O-17 is settled.
Images, built from 91c001c7d520e814c7bdba8bebddb371ef417077:
- Postgres:
hikarisystems/simmer:v0.5.0(linux/amd64, linux/arm64) - SQL Server:
hikarisystems/simmer:v0.5.0-mssql(linux/amd64)
Full Changelog: v0.4.0...v0.5.0
v0.4.0
v0.4.0
A feature release: header_rewrites, a per-route regex replacement over one
named header's value. Nothing else changes: a configuration that does not use it
behaves exactly as on v0.3.2.
Why. A List-Unsubscribe carrying a per-message token —
<https://www.meddoc.net/unsub.cfm?13323193_418550_3_9011119906_90535> — could
not be pointed at the link proxy. set_headers can only assign a template, and
body_rewrites only reaches text/* parts, so the body links went through the
proxy while the header, and a mailbox provider's one-click unsubscribe, did not.
identity:
header_rewrites:
- header: List-Unsubscribe
pattern: '<https://www\.meddoc\.net/'
replacement: '<https://link-pmps.healthcarematch.com/'- Order:
remove_headers, thenheader_rewrites, thenset_headers, so an
explicitset_headersvalue still wins. Templates still read the message as
it arrived. - Matched on the decoded value: folded lines are unfolded and RFC 2047
encoded-words decoded before any pattern runs. A header no rule changes
leaves byte for byte as it arrived. - Refused at startup: a rewrite of
From,SenderorMessage-ID(use
set_headers); rules that change their own output (not overridable by
unstable_headers); replacement text that could not form a valid header; and
a reference to a capture group the pattern does not have — including the
$1atrap, which the error explains. - New metric:
simmer_header_rewrite_skipped_total{route,header,reason},
for a header that could not be rewritten and went out unchanged — an unknown
charset, a broken encoded-word, raw 8-bit bytes, or a result too long to
fold. Worth alerting on if you rely on a rewrite.
Now in the specification. SPEC.md is amended for it: §4.1's example, three
§4.2 startup rules, §6.1 step 5a, §6.2 (the mechanism, its order and how it meets
§1.1's cutover invariant), §6.6 and §9.1. DECISIONS.md D-089 records the
reasoning, and open question O-16 is settled.
Images, built from 51da15b3f9f4cb617c02e77e1dcbe049d437b15d:
- Postgres:
hikarisystems/simmer:v0.4.0(linux/amd64, linux/arm64) - SQL Server:
hikarisystems/simmer:v0.4.0-mssql(linux/amd64)
Full Changelog: v0.3.2...v0.4.0
v0.3.2
v0.3.2
A documentation release. The server behaves exactly as v0.3.1 does — nothing
under src/ has changed between the two, so this is that binary plus a version
string and one comment in the shipped simmer.yaml. If you are on v0.3.1 there is
no functional reason to move; if you are on v0.3.0 or earlier, upgrade for v0.3.1's
fixes, which this release documents properly.
What the documentation now says that it did not:
-
The capture's flush policy. A record reaches the file within ten buffered
lines or 500 ms of quiet, sotail -fon the current bucket keeps up and a
replay of a range that has just ended finds its last records. These are flushes,
notfsyncs —on_error: deferremains the only mode that is durable before the
client is answered. README, DOCKERHUB and the shippedsimmer.yamlall say so. -
What to alert on, and what not to.
simmer_capture_disk_bytesis the right
gauge from v0.3.1: incremented as the writer flushes, recounted from the
directory by each retention sweep. On v0.3.0 and earlier it only moved on the
hourly sweep, so it read 0 for the whole first hour — measured over a 1-hour
soak, 0 for 59.9 of the 60 minutes while 904 MB accumulated. Anyone reading these
docs against an older image is now told to usesimmer_capture_bytes_total
instead, or upgrade, rather than trusting a number that was not moving. -
docs/STATE.md, re-snapshotted. It was taken at the end of phase 11 and
predated the link proxy, the SQL Server build, the capture,server replay, the
generalised soak tier and both releases. It now carries both feature sets'
verification numbers, the current module layout, the known defects — F17 and the
mssql migration race as fixed, the one-hour memory gate's resolution limit as
open — and which index serves which query on both backends, verified against the
planner rather than read off the DDL. -
The test tiers. The capture is not the soak's:
SIMMER_CAPTURE=onruns any
tier with it on, except the acceptance tier, which reads its config from the
image and so has no generated twin — and now says so instead of running
uncaptured.
Images, built from b4a6a447190ced9528258e610a30ece710b61bf8:
- Postgres:
hikarisystems/simmer:v0.3.2(linux/amd64, linux/arm64) - SQL Server:
hikarisystems/simmer:v0.3.2-mssql(linux/amd64)
Full Changelog: v0.3.1...v0.3.2
v0.3.1
v0.3.1
A patch release: no new configuration, no new metric, no new subcommand, no change
to any reply. What moves is behaviour inside things 0.3.0 already had.
The capture reaches the file promptly. It writes through a 256 KiB buffer, and
the only thing that emptied it on a live stream was a one-second tick — at the 4 KiB
end of a real message mix that is forty-odd records, so tail -f on the current
bucket showed nothing for long stretches. Now it flushes on ten buffered lines, or
500 ms of quiet, with the one-second tick kept as the backstop. Still a flush and
not an fsync: on_error: defer remains the only mode that puts a record on the
platter before the client is told anything.
simmer_capture_disk_bytes is live. It was written only by the retention
sweeper, on a one-hour interval whose first pass runs at startup against an empty
directory, so the one gauge documented for "will this fill the volume" read 0 for
the whole first hour — measured over a 1-hour soak, 0 for 59.9 of the 60 minutes
while 904 MB accumulated. The writer now adds what each flush pushes and each
sweeper pass still recounts the directory, so it is live between passes and
self-correcting across them. If you alerted on it, that alert now works.
Two replicas of the -mssql image can start at the same time. The migration
lock is now taken before simmer_migrations is created rather than after, so the
loser of that race no longer refuses to start with "There is already an object
named 'simmer_migrations' in the database".
Gates: 1,068 tests on the default build, 937 on the mssql build against SQL Server
Express, fmt and clippy clean on both feature sets. The T4 soak tier ran for an hour
against the mssql build on Express with the capture on — 36,010 messages an instance
with nothing deferred or refused — and is recorded in docs/SOAK.md §11.
Images, built from 3f667afb6f014763fd2708dab5ddd0df44586b48:
- Postgres:
hikarisystems/simmer:v0.3.1(linux/amd64, linux/arm64) - SQL Server:
hikarisystems/simmer:v0.3.1-mssql(linux/amd64)
Full Changelog: v0.3.0...v0.3.1