Releases: TikHub/Spinneret
Release list
v0.1.4
Four fixes found by running 0.1.3 rather than by reading it: an upgrade that
crash-looped for six seconds, a security opt-out that could not be set, an
update check that could not be turned off, and a proxy-trust setting that had
been living in one machine's working tree instead of the repository. No schema
migration and no configuration to change; upgrading is a pull and a restart.
Fixed
-
The notification SSRF opt-out could not be set on a Compose deployment. 0.1.3 turns the delivery
guard on for everyone and namesSPINNERET_NOTIFY_ALLOW_PRIVATE_TARGETSas the way back in, but the
variable was missing from the Compose environment block — and Compose delivers only the variables that
block names, so the documented escape hatch did not exist on the default deployment. Only the Compose
files changed, so the 0.1.3 image is unaffected:git pullin the deployment directory is enough. -
A restarting Valkey no longer takes the server down with it. An instance answers
LOADINGto every
command while it reads its AOF back after a restart, and the server treated that as a fatal startup
error. It exited, the supervisor restarted it, and that loop did eventually succeed — but the backoff
grows while the instance needs a fixed time to load, so the larger the dataset the further recovery
lags behind it. Startup now re-pings for up to two minutes while the reply isLOADING, and still
fails at once on anything else: a wrong password does not become right by waiting.The Compose health check hid this.
valkey-cli pingexits 0 even when the reply is an error, so the
container reported healthy while it was still refusing work anddepends_on: service_healthylet the
server start into it. It now matchesPONG. -
Disabling the update check is possible again. The checker reads an empty URL as "make no outbound
call" and.env.exampleoffered exactly that, but a blank value reads as unset and falls back to the
built-in feed, so the documented configuration was unreachable. The variable was also missing from the
Compose environment block, so the check could not be pointed at a mirror either. Both are passed
through now, and the switch is a separateSPINNERET_UPDATE_CHECK_ENABLED— it has to be, because
Compose delivers an empty string for every variable an operator has not set, so a blank-means-off rule
would disable the check on every deployment that left it alone. -
The deployment can say which proxies may set
X-Forwarded-For. Caddy overwrites the header for an
untrusted peer, which is right when the published port is the edge. Behind something that terminates
the connection first — a Cloudflare Tunnel, an nginx on the host, a cloud load balancer — Caddy's peer
is the Docker bridge gateway, and it replaced the real client address with172.18.0.1in every audit
entry and risk event.CADDY_TRUSTED_PROXIESnow carries the trusted networks, defaulting to loopback
only: no request in this topology comes from loopback, so out of the box nothing is trusted.
Full Changelog: v0.1.3...v0.1.4
v0.1.3
Two security fixes, both reachable only by an authenticated caller but both
across a trust boundary a multi-tenant deployment relies on. No schema
migration and no configuration to change; upgrading is a pull and a restart.
A deployment that deliberately sends notifications to internal hosts now has
to say so with SPINNERET_NOTIFY_ALLOW_PRIVATE_TARGETS=true.
Security
-
Webhook delivery could reach internal services and cloud instance metadata. Notification channels
(webhook, Feishu, DingTalk, WeCom, Telegram) validated only that the URL washttp(s)and then dialled
with the default transport. Nothing stopped a channel from targeting169.254.169.254or a service on
the private network, and becauseTestChannelanswers synchronously and the chat providers echo part
of the response body, the answer came back to whoever configured it. Creating a channel needs
notify:write, held by tenantadmin/owner— in a multi-tenant deployment those are customers, not
infrastructure operators, so this crossed a trust boundary.Delivery now refuses non-public targets in the dialer's
Controlhook, which runs after DNS
resolution and therefore also defeats DNS rebinding: a hostname that resolves to a public address when
the channel is saved and a private one when it is used is still blocked. Redirects were already
disabled. SetSPINNERET_NOTIFY_ALLOW_PRIVATE_TARGETS=truefor a trusted single-tenant deployment that
deliberately notifies internal hosts.The blocked set is wider than Go's own classification.
netip.Addr.IsPrivateis RFC 1918 and RFC 4193
only, so shared address space would still have gone through — and that is where Alibaba Cloud keeps
instance metadata, at100.100.100.200, rather than in the169.254.0.0/16every other provider uses.
Carrier-grade NAT,0.0.0.0/8,192.0.0.0/24,198.18.0.0/15and240.0.0.0/4are refused too, as are
the IPv6 forms that carry an IPv4 address inside an otherwise ordinary global-unicast address: NAT64
(64:ff9b::a00:1) and 6to4 (2002:0a00:0001::) both looked public while delivering to10.0.0.1.Thanks to @xuemian168 (#19).
-
One identity-import row could pin a CPU for hours. A row's
_tagsfield was de-duplicated with an
O(n²)linear scan and nothing capped the token count, so only the 32 MiB stream budget bounded the
work: 50 000 distinct tags already cost 1.6 s, and the growth is quadratic — a full-budget field runs for
hours on the parsing goroutine while the rest of the fleet shares that CPU. Import needs
identity:writeon the target site, which the node and worker tokens carry, making it the most widely
distributed credential that could trigger it.Tags and labels are now capped at 64 per row — the limit
UpdateIdentityRequesthas always enforced, so
this closes an inconsistency rather than tightening the API — de-duplicated inO(n), and the scan
stops at the cap instead of materializing every token.Thanks to @xuemian168 (#18).
Added
- Adding one identity no longer means hand-writing JSON. The import dialog takes JSON Lines or CSV, the
payload editor is a raw JSON editor, and there is noCreateIdentityRPC — so a single cookie had to be
escaped into JSON by hand. The identities page now has a New button whose form is built from the
selected type's fields: a textarea forcookie_mapandjson, a switch forbool, an input for the
rest, plus the_account,_regionand_tagsmetadata. It submits one JSONL row through
ImportIdentities, so normalization, deduplication and activation stay the rules a file import already
follows. Acookie_mapfield takes a pastedCookieheader as-is and decodes a browser export array.
Changed
- The console moved to React 19 and Vite 8. Vite 8 cuts the production build from roughly 7.5 s to under
a second. No API or configuration change. - Dependency updates:
@types/node,prettier,docker/setup-buildx-action,prometheus/client_model,
protovalidate, and the Python SDK'shatchling,respxandpytest-asynciodevelopment pins.
What's Changed
- fix(identity): prevent import tag/label CPU DoS (O(n^2) de-dup) by @xuemian168 in #18
- fix(notify): block SSRF to private targets in webhook delivery by @xuemian168 in #19
- chore(deps-dev): bump prettier from 3.9.7 to 3.9.8 in /web in the console-minor-and-patch group by @dependabot[bot] in #12
- chore(deps-dev): update hatchling requirement from >=1.32.0 to >=1.32.3 in /sdk/python by @dependabot[bot] in #10
- chore(deps): bump buf.build/gen/go/bufbuild/protovalidate/protocolbuffers/go from 1.36.12-20260825204119-511051f7f437.1 to 1.36.12-20260825204119-511051f7f437.2 by @dependabot[bot] in #7
- chore(deps): bump github.com/prometheus/client_model from 0.6.2 to 0.6.3 in the go-minor-and-patch group by @dependabot[bot] in #6
- ci(deps): bump docker/setup-buildx-action from 3 to 4 in the actions-all group by @dependabot[bot] in #11
- chore(deps): bump react, react-dom and @types/react in /web by @dependabot[bot] in #16
- chore(deps-dev): update respx requirement from >=0.22 to >=0.23.1 in /sdk/python by @dependabot[bot] in #9
- chore(deps-dev): update pytest-asyncio requirement from >=0.24 to >=1.4.0 in /sdk/python by @dependabot[bot] in #8
- chore(deps-dev): bump vite from 6.4.3 to 8.3.0 in /web by @dependabot[bot] in #14
- chore(deps-dev): bump @types/node from 22.20.3 to 26.6.2 in /web by @dependabot[bot] in #13
New Contributors
- @xuemian168 made their first contribution in #18
Full Changelog: v0.1.2...v0.1.3
v0.1.2
Two fixes, both for settings that looked like they worked and did not. No schema
migration, no configuration to change; upgrading is a pull and a restart.
Fixed
-
The console's ClickHouse retention setting now reaches ClickHouse. It stored a number and changed
nothing: the tables carry a TTL of their own, ClickHouse is the only thing that enforces it, and no
ALTER TABLE ... MODIFY TTLwas ever issued. A restart did not help either — startup migrated the
tables with the value from the environment and ignored the stored one. So the page reported 30 days
over tables that expired at 90, and the disk grew at the old rate until somebody looked.Changing it now alters the tables in the same call, and a failed
ALTERis reported instead of leaving
a stored number that means nothing. Startup resolves the setting before migrating, so a value set from
the console survives a restart. Clearing the setting applies whatever it falls back to. With ClickHouse
disabled the value is recorded for whenever one is connected. -
Nine documented environment variables could not be set on the Compose deployment at all. Compose
delivers only the variables itsenvironmentblock names, andSPINNERET_CLICKHOUSE_TTL_DAYS, the five
SPINNERET_RETENTION_*,SPINNERET_LATE_REPORT_WINDOW,SPINNERET_STREAM_MAXLENand
SPINNERET_MAX_WATCHERSwere not among them. Putting any of them in.envdid nothing — the server
kept its built-in default and nothing said otherwise. Found while setting a ClickHouse TTL on a real
deployment, where 30 days in.envleft the tables on 90.They are passed through with an empty default rather than the documented one, which matters: a variable
that always arrived with a value would count as explicitly set, pin its setting to the environment, and
make Settings → System show it read-only on every Compose deployment.TestABlankValueIsNotSetholds
that invariant down.
Full Changelog: v0.1.1...v0.1.2
v0.1.1
A patch release: no schema migration, no configuration you must change, and every default unchanged.
Upgrading is spnrctl pull && spnrctl up -d --wait. Two of the fixes below are things that were wrong
in 0.1.0 without saying so — container logs had no size limit at all, and the installer's default image
was one that only TikHub could pull.
Added
-
Retention is editable from the console, at Settings → System, instead of only through environment
variables and a restart. The five PostgreSQL retentions and the ClickHouse TTL are joined there by the
alert history window, which was hard-coded at 90 days and had no variable at all. A change applies on
the next hourly maintenance pass.The two sources compose in one direction: an environment variable that is explicitly set wins and
pins the setting, which the console shows read-only next to the name of the variable to remove. That
keeps the guarantee a deployment managed from a file depends on, and since nothing in the shipped
.envsets a retention variable, every existing deployment starts with all of them at their defaults
and nothing about its current behaviour changes. Values are stored as one row insystem_settings, which
has existed since the first migration and already holds the notification watermark, so there is no new
migration.Reading them needs only a console session — what a deployment keeps its data for is not secret, and an
operator looking at a filling disk should not need a role binding. Changing them requires a platform
administrator, because it is deployment-wide and destructive, and every change is audited, accepted or
refused: retention decides how long evidence is kept, so the change to it is evidence.Deliberately not editable here: anything encoded into data (
SPINNERET_REPORT_SHARDSlives in every
lease id), anything needed before the database can be reached (its own URL, the key-encryption keys),
and the acquire and report hot paths. The test for inclusion is not whether someone would like to
change a value but what happens if it changes between two reads. -
Every release is published to Docker Hub as well as to GitHub Packages. It is one build pushed to
both registries rather than a build each, so both serve the same digest and cannot drift apart, and
the four tags (v0.1.0,0.1.0,0.1, andlatestfor a non-pre-release) are identical on each.
The Docker Hub image istikhubio/spinneret, and it is the one to pull from a host that cannot read
GitHub Packages. Docker Hub needs theDOCKERHUB_USERNAMEandDOCKERHUB_TOKENsecrets; without them
the run publishes to GitHub Packages alone instead of failing, which is what a fork sees. The namespace
is that username unless the repository variableDOCKERHUB_REPOSITORYoverrides it.
Changed
- The installer gives Valkey one I/O thread on a host with two cores or fewer, instead of the
default four. Four buys tail latency on a machine with cores to spare and oversubscribes one that runs
PostgreSQL, ClickHouse, the server and the load balancer on the same two. Valkey executes every command
on its main thread whatever this is set to, so the cost is tail latency at high rates, not throughput.
The installer already scales the per-container memory ceilings to the host; this is the same idea for
the CPU. - The installer now pulls from Docker Hub by default,
tikhubio/spinneretinstead of
ghcr.io/tikhub/spinneret. GitHub Packages is private for this project, so the old default answered
401 for everyone outside the organisation and the installer quietly fell back to building from source —
which is the slow path the published image exists to avoid. The same build is on both registries at the
same digest, andSPINNERET_IMAGEstill points at GitHub Packages, a private mirror or a fork. - The
releaseworkflow can be run by hand against a tag that already exists, from Actions → release
→ Run workflow. It rebuilds from the tag and publishes the images without touching that tag's GitHub
release, which is how a registry added after a release was cut gets the images it missed. Itslatest
input exists so that publishing an older tag does not move:latestbackwards.
Fixed
-
Container logs are capped. Every service in the Compose files now sets
json-filewith
max-size=20mandmax-file=10. Docker's default keeps container output forever, so this was the one
store in the stack with no bound of any kind — not governed by a retention variable, not trimmed by a
job, and not in any of the three volumes an operator is told to watch. It writes to the Docker data
root, which is the filesystem PostgreSQL and ClickHouse are on. The retention table in the operations
runbook now lists it alongside the settings, and the disk alert covers the data root and not only the
volumes. -
Settings → System no longer tells a build made from source that it is on the latest release. A
build with no release number cannot be ordered against one — it may well be ahead of it — so the card
now says the build came from source and shows the latest release beside it without claiming either is
newer.CheckForUpdateResponsecarries the new fieldcurrent_is_releasefor the distinction;
update_availablewas already correct and is unchanged. Found by checking the console against the
real feed immediately after publishing v0.1.0, which is the first time the "unversioned build, a
release exists" combination could occur. -
The notification service reported itself ready before it had subscribed to the event bus, so an
event published in that gap reached no handler. The bus delivers to the handlers a channel holds at the
moment of publication and keeps no backlog, so such an event is lost rather than late.Service.Run
setsrunningon its first statement because that flag is the guard against a secondRun, several
statements before the subscription exists; readiness is now the separateService.Subscribed(), set
once the delivery workers and the bus handler are in place. Anything that publishes and then waits for
the result should wait on that. Two CI builds failed on this, each in whichever notify bus test lost
the race. -
TestShedAfterRedisReplyKeepsExhaustedno longer depends on the order the scheduler's tests run in.
It asserts that only the first acquire attempt reaches Redis by counting commands, andScript.Exec
sendsEVALSHAand falls back toEVALwhen the server answersNOSCRIPT, so the first attempt of a
cold run costs two commands instead of one. The test now loads the script before it starts counting.
Full Changelog: v0.1.0...v0.1.1
v0.1.0
First release, open source under the Apache License 2.0 and maintained by
TikHub. Spinneret is a control plane for fleets that share scarce,
rate-limited credentials and egress: it leases identities and proxies to nodes, turns the request
results they report into cooldowns, bans, health scores and circuit breaking, and distributes versioned
configuration and secrets. Nodes need only a server URL and an API token.
Added
Core path — identity scheduling and reporting
LeaseService(Acquire,AcquireBatch,Renew,Release): candidate sampling, availability filtering
and lease writing in a single atomic Redis Lua script, with no PostgreSQL round trip on the hot path.- Rotation policies: strategies
weighted_random,least_recently_used,round_robin,best_health;
lease TTL and lifetime cap, concurrency limit, reuse interval withacquired/releasedanchor and
endpoint-group/site scope, per-window quotas, sticky sessions, warm-up and probe weighting. - Identity types with typed payload fields (
string,number,bool,cookie_map,json,secret_ref),
JSON-Schema-validated imports (JSON Lines and CSV, up to 50 000 rows per call, dry run), deduplication by
unique_by, payload versioning and a six-segment delivery rendering that hides type details from nodes. ReportService/Report: batches of up to 500 reports, per-report validation, idempotency byreport_id,
asynchronous ingest into sharded Redis streams.
Admission control on the Acquire path
- Acquire concurrency at Redis is bounded per instance, so offered load above the knee becomes shedding
instead of congestion collapse.SPINNERET_ACQUIRE_FLEET_INFLIGHT(default64, range 0–65536) is the
budget for the whole fleet, and each API instance admits that number divided by the live API instances
it sees in a Redis heartbeat set, clamped to[4, 4096].SPINNERET_ACQUIRE_MAX_INFLIGHT(default0,
max 4096) pins one instance's limit instead and stops its heartbeat — pin it on every API instance or on
none.SPINNERET_ACQUIRE_FLEET_INFLIGHT=0turns admission control off entirely, which is also the A/B
lever for measuring it on one build. Release, renew, reap and ingest are deliberately left ungated: that
asymmetry is the mechanism, because it bounds the queue arelease.luawaits behind. - An attempt waits up to 50 ms — never longer than its own remaining
wait_ms— for a permit, and the time
it spent waiting is credited against its retry rung, so a givenwait_msbuys the same number of attempts
it would without the gate. What cannot be admitted is shed with the reasonoverloaded
(unavailable, HTTP 503) and a retry hint jittered into 100–200 ms. A shed issues zero Redis
commands, which is why retrying it is always safe and why it cannot deepen the queue.AcquireBatch
takescountpermits, so a 50-lease batch is charged for the Lua work it really does. - Two failure modes, two statistics results.
exhaustedmeans the identity pool is empty — add identities
or widen the rotation policy;overloadedmeans the instance is at its acquire concurrency limit — add
capacity or offer less load. Alert on each for its own cause and excludeoverloadedfrom
acquire-failure alerts, since shedding is the healthy response to overload. The console's acquire failure
ratio counts both. - Metrics:
spinneret_acquire_script_seconds(theacquire.luaround trip alone, sharing buckets with
spinneret_acquire_duration_seconds),spinneret_acquire_admission_total{result}(immediate,queued,
shed_no_wait,shed_queue_full,shed_timeout,shed_canceled),
spinneret_acquire_admission_wait_seconds, and the gaugesspinneret_acquire_inflight,_queued,
_inflight_limit,_peers,_peer_beat_age_secondsand_peer_beat_failures_total. The gauges exist
only while admission control is on;_peersis also absent when the limit is pinned.
Risk-control loop
- Signal policies: configurable rules over HTTP status, business code, error kind, markers, URI, method,
latency and size, producing twelve outcomes; optional trust of node-proposed outcomes. - Action policies: cooldown, expire, quarantine and ban at identity-endpoint, identity-site, identity,
account, proxy-site and proxy scope; exponential backoff with caps, escalation ladders, failure-streak
resets and a shadow mode that records without enforcing. - Health scoring: EWMA with time decay towards a baseline, per-endpoint low-score cooldowns and automatic
quarantine; identity lifecyclepending → active → quarantined / banned / expired / disabled / retired. - Cross attribution between identities and proxies, so a bad proxy is not charged to the identities that
used it. - Circuit breakers per endpoint group: sliding window with buckets, three states with probe leases, manual
open and close, and optional revert of the cooldowns applied in the tripping window. RevertActionsfor bulk rollback of automatic actions by time range, site, policy, rule and action kind,
with dry run and optional health/failure resets.
Infrastructure
- Config center: versioned items with drafts, publish, rollback and diffs;
WatchConfiglong polling;
${secret:path}references resolved at delivery; reserved read-only_runtimegroup exposing breakers
and site switches. - Vault: AES-256-GCM envelope encryption (KEK → DEK → data) with file and environment KEK providers,
online KEK rotation and a resumable rewrap job, secret versions, expiry alerts and audited reads. - Proxy pool: kinds, regions, providers, tags, concurrency limits and session templates; assignment modes
none,pool,bind_identityandregion_match; periodic health checks with exit-IP and GeoIP
enrichment; imports from URL lines, JSON Lines and CSV. - Tenancy and access control: tenants → namespaces → sites, console roles
viewer/operator/admin/
ownerwith per-namespace and per-site bindings and extra permissions; node API tokens with scoped
grants, IP allow-lists, rate limits and expiry; Argon2id passwords, login throttling, sessions and CSRF
protection. Token names are unique among usable tokens only, so revoking a token frees its name for the
replacement. - Hot-state rebuild from PostgreSQL (
spnr rebuild), full or per site, with periodic snapshots of health
state and automatic rebuild when Redis has lost its data. - Report stream retention: each shard owner trims its Redis stream to the consumer group position
(XTRIM … MINID ~, floored by the oldest pending entry), so a drained stream does not keep every
processed entry untilSPINNERET_STREAM_MAXLENevicts it. - Observability:
/healthz,/readyz(PostgreSQL, Redis, hot state, catalog, draining), Prometheus
/metrics, optional OTLP tracing and an optional ClickHouse-backed request explorer. An optional
profiling listener,SPINNERET_PPROF_ADDR, servesnet/http/pprofon its own address; it is off by
default and is never mounted on the API or metrics listener.
Console and delivery
- Embedded React console covering every module across 22 routes: overview, heatmap, identities and
identity types, accounts, proxies, sites, policies, breakers, config, secrets, request explorer, risk
events, notifications, tokens, users, audit log, tenants and system. Bilingual (English and Chinese),
light and dark, live updates over SSE. - Settings → System reports the running build and checks the published releases on demand. Nothing polls:
the deployment reaches out only when an operator presses the button, the answer is cached for an hour,
andSPINNERET_UPDATE_CHECK_URL=""disables the check entirely. The check runs server-side, so the
console'sconnect-src 'self'CSP stays as it is and the feature works where the browser has no route
out but the host does. - Notifications: webhook (HMAC-SHA256 signed), Feishu, DingTalk, WeCom and Telegram channels; eleven alert
kinds with de-duplication, per-site routing, test deliveries and delivery history. spnrCLI:migrate,admin init,token create,rebuild,kek generate|status|rewrap,seed,
config check,healthcheck,version.- Python SDK (sync and async, httpx + pydantic v2) and Go SDK (Connect JSON or gRPC,
github.com/TikHub/Spinneret/sdk/go), both with lease helpers, batching reporters, config watchers with
local snapshots and typed errors. - FastAPI example crawler (
examples/fastapi-crawler) with a one-command quickstart, and a mock target
site and proxy for testing without touching a real platform.
Install, deployment and documentation
- A guided one-command installer,
install/install.shwith an identical Chinese twin
install/install.zh.sh. Docker only: it detects the host, offers to install Docker, clones the
repository, generates every password and the key-encryption key on the machine, writes the Compose
overrides for this host, pulls the published image or falls back to building from source, runs the
migrations, waits for/readyz, creates the first administrator, and writes aspnrctlcontrol script.
Re-running it opens a management menu: status, upgrade, accounts, tokens, hot-state rebuild,
configuration, health, logs, restart, backup, restore, disk and uninstall.install/README.mddocuments
every question, flag and file it writes. - Docker Compose deployment (
deploy/compose) with PostgreSQL, Valkey, ClickHouse, migrations, two server
replicas and a Caddy load balancer with readiness-based health checks, plus the profilesinit,
observability,test,loadtestandexample. - A complete bilingual manual in
documents/: 21 pages in English and Simplified Chinese, from a quick
start to a performance and tuning guide, with an index indocuments/README.md. Every page was verified
against the source.
...