Repository navigation
26.8.17-1
Changelog
⚙️ CI/CD
- Rate-limit the unauthenticated surfaces at the edge
/oauth/authorize and /oauth/token are reachable without a credential and
both fetch a caller-supplied Client ID Metadata Document before anything is
validated. Traefik middleware rather than src/rate_limit.rs so the spam never
reaches the pods.
- Stop a scheduled pipeline firing the push-only jobs
Nothing in this repo creates a pipeline schedule today, but anyone can add one
from the GitLab UI in about ten seconds, and the result would be surprising.
⚙️ Miscellaneous Tasks
- Refresh the openapi artifact
node_job_stats landed the NodeJobStats/NodeJobStatsHistoryEntry schemas
and a UserNode change without regenerating openapi.json, leaving the
openapi_artifact_is_current test red.
- Add claude design-sync config for the app design tokens
Tokens-only sync of the Tauri frontend design language (Tailwind v4 theme,
glass/purple palette, component classes) to the "Bitping Node" claude.ai/design
project. Sync inputs: converter config, full token set, design-context CSS
resets, empty bundle entry stub, conventions header, and re-sync notes.
- Keep fuzzing a local, on-demand workflow
Nothing runs the fuzz targets automatically — they are a thing you reach for
while working on the wire or auth code, via just fuzz-list / fuzz /
fuzz-for.
Two fixes to make that actually work:
The recipes passed fuzz/seeds/<target> as the only corpus directory. libFuzzer
writes newly interesting inputs into the FIRST directory it is given, so a local
run would have written into the committed seeds. The gitignored
fuzz/corpus/<target> now goes first and the seeds stay read-only.
The nix devshell pins stable and ships neither rustup nor cargo-fuzz, so every
recipe ran head-first into a confusing toolchain error. A _fuzz-preflight
dependency checks for both up front and prints what to install, exiting non-zero
before any build starts. It installs nothing.
$ just fuzz-list
fuzzing needs rustup — this shell's toolchain can't build fuzz targets.
install rustup: https://rustup.rs
then: rustup toolchain install nightly
fuzzing needs cargo-fuzz.
install it: cargo install cargo-fuzz
Run these outside the nix devshell. See docs/fuzzing.md.
docs/fuzzing.md drops its CI chapter for the corpus-directory rule and local
crash triage: reproduce, cargo fuzz tmin, then land the minimised input as a
named unit test next to the code — which is how the short-base58 panic earlier
in this branch ended up with a regression test rather than a crash file.
⛰️ Features
- Add MCP server exposing network data, job dispatch, and usage tools
Streamable-HTTP MCP server at {route_prefix}/mcp on rmcp 3.1 (MCP spec
revision 2026-07-28), authenticated with the existing customer API key.
Public node/network tools need no auth; jobs_perform_* and my_api_usage
require x-api-key from an active BITPING_API_ACCESS subscription.
Served statelessly for every protocol version: legacy_session_mode(false)
extends rmcp's 2026-07-28 behaviour to older clients, so no session lives in
process memory. The deployment runs two replicas with no session affinity, so
a pod-pinned session would break as soon as a request reached the other.
Not part of the utoipa spec — it is a nested tower service, so it never
appears in openapi_spec() or ROUTE_MANIFEST.
- OAuth 2.1 authorization server for the MCP endpoint
An MCP client can obtain its own API key from a logged-in dashboard session
instead of the user pasting one. oxide-auth owns the grant (PKCE S256,
single-use codes); we supply the four primitives plus the MCP-specific
surface it has no notion of — RFC 9728/8414 discovery and Client ID Metadata
Documents. Every issued token is an ordinary api_keys row, so there is no
parallel credential store and no second revocation path.
Client identity is CIMD only. MCP 2026-07-28 deprecates RFC 7591 dynamic
registration, and Claude Code has preferred a hosted metadata document since
v2.1.81, so a registration endpoint would be an unauthenticated write path
with no remaining users. A client_id is an https URL the client hosts; we
fetch it behind SSRF fencing (no redirects, pinned DNS, private-range
refusal, capped body) and store nothing.
Requires bitping-db !4: oauth_authorization_code, api_keys.oauth_client_id,
and the drop of oauth_client.
- Dashboard rows for the MCP tools and OAuth provider
A JSON-RPC error rides inside an HTTP 200, so the existing status panels
cannot see a failing tool call. The OAuth row is the only visibility into a
flow that has no other operator surface: client-document fetches, token
issuance, and the retention sweep that runs off the issue path.
- Publish the MCP endpoint at bitping.com/mcp
Customers get a URL with no api subdomain and no version in it. Traefik routes
bitping.com/mcp to this service and rewrites it onto the existing /v2/mcp
mount, so nothing moves server-side.
MCP discovery is origin-anchored, so a path route alone is not enough: RFC 9728
puts the protected-resource document at the RESOURCE's origin root, and the
resource value plus the 401 challenge were both derived from public_base_url.
mcp_public_url now anchors all three, and well_known_router mounts whatever
path the advertised document claims — the api-host form keeps working, so a
client already pointed at api.bitping.com/v2/mcp does not break.
- Meter every consent outcome
access_denied is all the OAuth client is ever told, and no pod logs reach the
Grafana Loki, so a failing authorization is currently undiagnosable from the
outside — the first real Claude Code attempt could only be narrowed by probing
production endpoint by endpoint.
bitping_api_oauth_consent_total labels all nine exits from check_consent, so
the branch a denial took is visible: denied_origin, denied_no_decision,
denied_by_user, denied_bad_owner, denied_bad_redirect, refused_key_limit,
refused_email_unverified, refused_unavailable, bounced_login, authorized.
Adds the matching dashboard panel and richer logs on the paths that had none.
- Return a refused approval to the consent page
A refusal used to reach the user as OwnerConsent::Denied, which oxide-auth
reports to the client as access_denied at its redirect_uri. For a native MCP
client that is a loopback port which is no longer listening, so the user saw
nothing: no success, no failure, no reason. That is exactly what the Astro
form regression looked like from the outside.
Only a deliberate Deny is the client's business now — RFC 6749 says tell it.
Everything we refuse ourselves (unreadable POST, unexpected origin, an owner
id that will not parse) goes back to consent with a reason the page renders.
The unreadable-POST guard moves ahead of the session check, alongside the
origin guard: both judge the request itself, neither needs to know who sent
it, and a malformed POST no longer costs a database round trip.
- POST /user/auth/logout revokes the presented token
Logging out only cleared a cookie; the JWT stayed valid for its full 7-day
expiry with no way for anyone to retire it. The only levers were a password
change, which nobody would guess means sign out, or rotating
ACCOUNT_HMAC_KEY, which signs out every user on the platform.
Each token now carries a jti, and logout writes that one id to revoked_token.
Per TOKEN, not per user: signing out on a laptop leaves the same account
signed in on a phone. The check rides the existing verify_user_hmac query
rather than a second round trip.
A token with no jti is refused outright. Honouring one would be a permanent
bypass — anything already leaked would keep working for its full expiry with
nothing able to retire it — so every session minted before this ships is
invalid and its owner logs in again once.
Requires bitping-db nick/revoked-token-list.
- Add business overview dashboard
Postgres + VictoriaMetrics dashboard covering revenue (active
subscribers by tier, base guaranteed MRR from verified Stripe
pricing), customer acquisition (signups/paying-customer conversion
per week), spend (node-operator earnings accrued vs. actual payouts
paid, via the same vw_latest_currency_price join used by the
payouts route), and a gross-margin proxy, plus network activity
context from prom-bitping. Every query dry-run validated against the
live bitpingdb-pgsql/prom-bitping Grafana datasources before being
committed.
- Add true revenue, churn and cost-attribution panels
Base MRR counted only flat subscription fees, so the legacy
Pay-As-You-Go tier — which has no flat fee and is the majority of
active accounts — contributed $0 to the headline revenue figure
regardless of consumption. Margin, derived from it, read negative
as a result.
Adds Total Revenue (flat + metered overage, computed per subscriber
over their own Stripe billing window using api-usage-reporter's unit
math) and True Gross Margin alongside the existing floor figures,
which are retitled to say what they actually are.
Adds a "Where The Money Goes" row explaining the economics: a
revenue-to-gross-profit bridge, the API-calls-vs-bandwidth cost
split, per-unit economics, per-tier list margins, and a text panel
walking through why the two figures differ. Adds a Retention row
(weekly cancellations, 90d churn rate).
All 9 new SQL queries executed against bitpingdb-pgsql and the full
27-panel dashboard dry-run validated against the Grafana v2beta1 API
before commit.
Also documents the new repo rule in CLAUDE.md and AGENTS.md:
dashboard and documentation changes may land directly on master
without an MR, provided the commit carries no code changes and the
docs do not reference unmerged code.
- Break node payouts down by feeder vs standard rate
Cost per billable unit sat above the $0.04 standard node rate and the
dashboard could not say why — feeder rate or job fan-out. Resolved:
~94% of payout dollars go out at the 2x feeder rate, and fan-out is
ruled out (node-side jobs / customer-side calls is ~0.66, below 1).
is_feeder is keyed off the CUSTOMER, not the node — an allowlist of
internal feeder-account user IDs in FEEDER_USER_IDS
(node-earnings-engine/src/main.rs:114-119). So the bulk of node spend
buys traffic for feeder accounts rather than serving paying customers,
and does not scale down with customer demand. That is the single
largest line in the business and it was previously invisible.
Adds Feeder Share of Node Payouts, Node Payouts by Rate (30d) and
Daily Node Payout Run-Rate by Rate, and rewrites the two panel
descriptions that carried the now-resolved fan-out caveat.
Also notes the any-feeder-wins rate-selection flaw on the bandwidth
leg (BIT-682) inline, since it makes the feeder bandwidth share an
overstatement.
Queries validated against prom-bitping and the full 30-panel dashboard
dry-run validated against the Grafana v2beta1 API before commit.
- Accept an API key on GET /user/nodes/
Adds two axum extractors. Auth is JWT-only and yields the full Claims.
AnyAuth additionally accepts an API key and yields only the resolved user
id; GET /user/nodes/ opts into it, so a customer can list their own nodes
with the same credential the jobs surface takes.
Auth is deliberately NOT widened to accept x-api-key. Doing so would hand
every API key the power to delete accounts, request payouts and rotate keys,
so accepting a key stays opt-in per route.
The key path reuses api_key_authenticate, so the 402 subscription gate
applies unchanged. A bearer token is routed by shape: three dot-separated
segments take the JWT path, anything else is treated as an OAuth-issued API
key. Deciding by shape rather than retrying a failed JWT keeps every
published JWT error reaching a caller that sent a real JWT.
Paging lifts the querystring validation into an extractor declared ahead of
AnyAuth, because axum runs FromRequestParts left to right — that is what
keeps the published 400-before-401 ordering, and a test pins it.
- Exchange an API key for a node PASETO
New ApiKeyAuthenticate RPC modelled on JWTUpgradeAuthenticate: verifies an
ed25519 signature over SHA256(api key bytes), honours api_keys.expires_at,
rejects a pubkey already bound to a different account, registers the pubkey
in user_keys, and returns a node token plus the owner's email so the node
can sign SHA256(email) on subsequent SignatureAuthenticate logins.
- Log in with an API key
bitpingd login --api-key and a BITPING_API_KEY rung in the startup
recovery ladder (after BITPING_EMAIL/BITPING_PASSWORD) exchange the key
over the new ApiKeyAuthenticate RPC. One-shot bootstrap: the exchange
registers the node's keypair, so every later start-up re-authenticates
via SignatureAuthenticate and the key itself is never stored.
- BIT-694 reputation persistence, weighted selection, hub-ward dispatch
Squashed feature implementation for BIT-694 phases 0-4: account-keyed
reputation (credibility/verdict rules in economics::, separated from
the billing formula which stays customer-claim-anchored per !113),
persisted + cross-hub-synced reputation store, reputation-weighted +
grief-pit-homophily node selection in connection_pool, and the
Phase-3 move of customer job-dispatch auth + billing hub-ward
(bitping-api's jobs routes become a thin proxy onto the hub's gRPC
public_jobs endpoint).
This was previously bundled into the Gherkin-conversion commit by
mistake — split out so the history actually reflects feature work
vs. test work, as it claims to.
🐛 Bug Fixes
- Pass ip_addr to poll_device_login in the device-code test
9e0afd2c1f added the ip_addr parameter to authenticator-service's
poll_device_login but missed this cross-service caller, so
cargo test -p bitping-api did not compile on master.
- Let a loopback OAuth client choose its port
RFC 8252 §7.3 requires the authorization server to allow any port on a
loopback redirect: a native client takes whatever ephemeral port the OS gives
it at request time. Claude Code registers http://localhost/callback in its
metadata document and asks for http://localhost:51219/callback, so the
verbatim bind refused it and every native MCP client failed authorization with
"invalid authorization request".
The port licence is loopback-only — off loopback another port is another
service, and matching across ports would let a client aim the code at it. The
bound value is now the REQUESTED url, since that is the port the code is
actually delivered to.
- A same-origin consent POST carries no Origin header
The consent page posts through the dashboard's own /api/v2 proxy, so the
approval is same-origin — and Safari omits Origin entirely on a same-origin
form submit. Refusing a missing header denied every approval from that
browser with access_denied, which is what the first real Claude Code
authorization hit.
A browser always attaches Origin to a CROSS-origin POST, so a missing one can
only be same-origin. The sibling-subdomain forgery this guards against — the
token cookie is scoped to .bitping.com, so docs.bitping.com could host the
form — is cross-origin by construction and still carries a header to refuse.
A literal null origin (sandboxed iframe) stays refused.
- Restore the same-origin consent fix reverted by !102
!101 taught the origin guard that a missing Origin header is a same-origin
POST, which is what the consent page performs — it posts through the
dashboard's own /api/v2 proxy, and Firefox and Safari both omit the header
there. !102 then restored solicitor.rs wholesale from an editor snapshot that
predated the fix, reverting both the guard and the table case covering it, so
CI stayed green while production went back to denying every approval.
bitping_api_oauth_consent_total, added by that same MR, is what caught it:
two probes with no Origin header both landed on denied_origin.
- Allow the published MCP host past the rebinding guard
rmcp's DNS-rebinding guard 403s any request whose Host is not on its
allowlist, and the allowlist was built from PUBLIC_BASE_URL alone. Publishing
the endpoint at bitping.com/mcp meant every real request arrived with
Host: bitping.com and was refused by the transport before auth ran — so a
client that had completed OAuth still could not call a tool.
The allowlist now covers both settings: the endpoint is published on one host
and mounted under another, and the Host header carries whichever the client
was handed.
- Return A/AAAA/CNAME/PTR records from DNS jobs
The node resolves these record types and the proto carries them as
DNSLookupResponse fields 9-12, but DnsResult never declared them and its
From impl never read them, so a job with lookupTypes ["A","AAAA"] came
back successful with every field empty.
Correct the MCP tool doc alongside it: it advertised "A"/"AAAA" as the
example values, which was the exact combination that returned nothing.
- Close the parity gaps the auth-extractor review found
Three of these are real regressions the extractor conversion introduced.
The token cookie could authenticate as an API key. token_from_headers
falls back to the cookie, so a dotless value there reached the key path — a
second, unintended credential channel onto every widened route. The key
channel now reads Authorization only, and a cookie can never carry a key.
An x-api-key preempted a valid JWT, so a caller sending both (a dashboard
session plus a stale key) started failing a request that used to succeed. A
dotted credential now outranks a key, and everything the key channel does not
claim falls through to authenticate untouched — which is what makes the
widening incapable of regressing a request that authenticated before.
A truncated bearer like aaa.bbb returned Invalid API key instead of the
pinned jwt malformed. Dispatch is on "contains a dot" rather than "has three
segments", so every dotted token keeps the full JWT error set.
Undecodable path segments turned 400 into 500 in CountryCode/ReportId,
and into 401 on the five handlers where Auth was declared ahead of Path.
The validation extractors now reject with a Response so axum's own rejection
survives byte-for-byte, and Path is ordered ahead of Auth everywhere.
- Resolve the full user from an API key, and 400 a bad id
The API-key lookup now selects email and name alongside the id, so both
strategies resolve the same User and JwtUser is a User plus the session
a key cannot prove. The .sqlx cache was regenerated for the changed query
with just bitping-api-sqlx-regen; exactly one entry moved.
A path segment that is not a uuid is a bad parameter, not a missing resource:
answering PASSKEY_NOT_FOUND / IDENTITY_NOT_FOUND claimed we looked one up
when nothing could have matched. PathUuid now rejects with 400
INVALID_PARAMETER before the handler runs, and /apikeys says the same for
the id it reads from the body.
CLAUDE.md records how to regenerate the cache, including the two ways the
recipe fails from an agent's shell: it needs nix develop --command for
protoc, and an absolute bitping-db path when run from a worktree.
- Pay cashouts only to the stored payout address
Cashout honoured a caller-supplied cashout_address, so any holder of a node
token (including one bootstrapped from a leaked API key) could send an
account's balance to an arbitrary destination. It now always resolves the
account's stored blockchain_account; the request's address/currency fields are
deprecated and ignored. Changing the payout address stays a dashboard-only
action, so a node key can trigger a payout but never redirect it.
The desktop payout modal reflects this: the destination is shown read-only
(managed in the dashboard) with an empty-state prompt when none is set, instead
of an editable field whose input the backend now ignores.
- Verify the federated API-token signature and honour expiry
federated_api_token_authenticate declared a signature field and never
verified it, embedding the caller's public key into the token claims on
trust; it now requires an ed25519 signature over SHA256(api_token bytes),
matching what p2proxy already sends. The key lookup also filters out
api_keys rows past their expires_at, which no auth path enforced.
- Standardise on "withdrawal" wording and i18n the payout modal
User-facing copy mixed "withdrawal" and "payout" (the modal said "Withdrawal
Address" but the hint said "payout address"), and three strings were hardcoded
in English so they stayed English under non-English locales. Standardise all
visible copy on "withdrawal" (the dominant existing term) and route the button,
address hints, empty state, and loading label through i18n across all eight
locales.
Internal identifiers (PayoutsService/Cashout gRPC contract, the payout DB
table, PayoutModal, the request_payout key) are wire/schema/internal names and
are intentionally left unchanged.
-
Reject a short base58 string instead of panicking
-
Stop logging the account password and MFA code
The desktop login command logged the cleartext account password and MFA code at INFO. The default log filter floor is info and SqliteLogger persists everything above DEBUG, so both were written to the node's unencrypted SQLite log DB and then readable via fetch_logs / exportable via export_logs. Log a credential-free breadcrumb instead.
BIT-691
- Remove the SSRF-prone favicon proxy
GET /misc/favicon.ico was a public, unauthenticated proxy that fetched arbitrary caller-supplied URLs server-side and returned the body — a read-SSRF into internal services and cloud metadata. It is unused: the dashboard app resolves favicons via its own endpoint that delegates to Google/DuckDuckGo (never fetches the caller's host from our origin, so no SSRF), and the marketing site is moving to the same shared resolver (website-monorepo !28). Delete the route, its service, the AppState field, the test double, and the openapi entry rather than harden a legacy proxy. openapi.json regenerated via --dump-openapi.
DEPLOY ORDER: website-monorepo !28 (marketing off bitping.com/img/favicon.ico) must land and deploy BEFORE this, or the homepage customer-logo favicons break.
BIT-690
- Bound node bandwidth credit to the corroborated minimum on disagreement
On a ByteCountMismatch/Fraud, clearing credited the node (customer+node)/4 per direction — strictly increasing in the node's own reported byte counts, which are unsigned self-claims (BandwidthReport signs only the session envelope, never the counts). A node could over-report to mint effectively unbounded earnings, paid downstream by node-earnings-engine, with reputation the only (uninenforced, metrics-only) penalty. Credit the corroborated min(customer, node) per direction instead, so over-reporting can never earn more than the honest counterparty conceded. The customer bill stays at 2*max.
BIT-689
- Verify the Postgres TLS certificate
The Postgres connection installed a NoVerifier that accepted any server certificate. Verify against the platform trust store (rustls-native-certs) instead, so a rogue or self-signed cert fails the handshake. The deploy must ensure the CNPG CA is trusted in the container's trust store, or the connection will be refused.
BIT-692
- Bind the SOCKS5 listener to loopback by default
The SOCKS5 listener hardcoded 0.0.0.0 and docker-compose published 1080:1080 to all host interfaces, so the shipped config exposed an unauthenticated, token-billed proxy to the LAN/internet. Add a host server-config field defaulting to 127.0.0.1, thread it into the listener bind, publish the docker port to 127.0.0.1 only, and set host: 0.0.0.0 in the docker Config.yaml (safe because the port is host-loopback-only, and the container needs 0.0.0.0 for docker's port forward to reach it).
BIT-693
- Trust the CloudNativePG CA for Postgres TLS
The rustls client built its trust store from the system CA bundle only,
so it refused the CNPG-signed Postgres server cert — CNPG signs from a
private per-cluster CA that is absent from the system trust store. Load
an optional database.ca_cert_path PEM into the RootCertStore alongside
the native roots, failing loud if the configured CA is unreadable or
empty (never falling back to accepting any certificate), and mount the
CNPG <cluster>-ca secret into the cronjob at /app/certs.
- Saturate the punitive bandwidth bill so a claim cannot wrap it to zero
Byte counts on a bandwidth report are peer-supplied and unsigned, and the
punitive arm doubles them. A customer reporting 2^63 outgoing bytes makes
2 * max(c, n) wrap to exactly zero on a release build, which bills that
direction at nothing — the same "submit garbage to escape billing" vector the
punitive rule was written to close, reached through arithmetic instead of
hashes. Saturate the doubling and the credit sum; the accumulator downstream
already saturates.
Cover the credit bound the same way: node credit must pin to the customer's
view and not move when the node inflates its own claim, and fall to the node's
own figure when it under-reports.
-
Order the signed job-result and peer-discovery collections
-
Verify every signed payload a peer sends before acting on it
Federation authenticates the connection, not the answer. Four consumers
took an Auth<T> off the wire and read .item without ever checking the
signature, so a peer whose token validated could put anything in the
payload:
bandwidth_report/ingest.rsbilled and credited fromauth.token
alone. Both PASETOs leak to the counterparty during a proxy session —
the customer's rides in cleartext inside the SessionInit payload every
exit node reads, and the node's rides in the ack — so either side could
attach a harvested token and have a stranger's account billed for its
traffic or credited for its bytes.query/fanout.rsandquery/resolve.rsmergedFoundNodesfrom peer
hubs, letting one hub inject arbitrary peer ids and addresses into node
selection.handshake/mod.rstreated an unverifiedHasPeerResultas grounds to
Backoff a node, so a peer hub could evict nodes from the mesh at will.
Rather than add four .verify() calls that a fifth consumer can forget,
make it unrepresentable. Unverified<T> is what inbound protocol types
now deserialise into: private field, payload reachable only through a
verify_* method, each returning Verified<T> — which itself has no
public constructor, no From, no Deref. Reading .item off the wire
no longer compiles. Both wrappers are #[serde(transparent)], so no wire
byte changes.
p2p/federated.rs is the one place a peer hub's response is opened.
Fan-out callers go through verified_or_dropped, which drops a bad
response with a warn plus federated_response_verification{phase,outcome}
so a misbehaving peer is attributable — one bad hub degrades to the same
outcome as a slow one rather than failing the query.
Verified<BandwidthReport> is threaded into the tracker: the byte counts
that move money can only enter the billing path with the proof attached.
Elsewhere the payload is consumed on the next line, so it is unwrapped at
the check.
- Bind a node's handshake token to the key that signed it
validate_handshake proved two things separately and never joined them: the
envelope signature verified against the pubkey in the request body, and the
PASETO validated against the trust root. Nothing required the two keys to be
the same, so the token was a bearer credential — present a harvested one, sign
with your own key, and NodeInfo.user_id came from the victim's sealed footer.
Their account is then credited for the attacker's node's work, and the attacker
inherits the victim's ip_hmac and location for pool placement.
Node tokens are harvestable by design: every node sends its own token to the
customer in the proxy session ack (p2p-bandwidth-protocol server_accept),
so any customer that has proxied through a node holds that node's credential.
The signature check already proves possession of the body key, so binding is
one assertion: the key the token claims must be that same key.
Failure mode is Reauthenticate, not a hard reject. Every mint path —
email/password, signature reauth, API-key bootstrap, device login, and the
local-auth-stub — sets the public_key claim to the key the caller itself
supplied and proved possession of, and UserSession persists the keypair and
the token in one record, so a mismatch means a stale pairing rather than a
different account. A fresh token resolves it. The one path that would not
self-resolve is an in-process identity change (the swarm caches my_key at
supervisor start while the token is read live), which is bounded by the
supervisor's restart loop; the reauth-tracker throttle keeps that case backing
off rather than spinning.
handshake_pubkey_binding{outcome=bound|mismatch|claim_undecodable} carries
the rate with its denominator, so the mismatch fraction is measurable in
production. Recommend watching it settle at zero before anyone tightens this
to a hard reject — the metric is the evidence for that decision, and this
commit deliberately does not make it.
- Refuse a handshake whose token does not bind the signing key
Answering Reauthenticate on a token/key mismatch invited the forger to
re-present the same stolen token. Every node hands its token to the customer in
the proxy session ack, so anyone who has served a session could sign with their
own key, present a harvested token, and have the victim's user_id — and the
victim's earnings — attributed to their node. Refuse it outright with a typed
TokenKeyMismatch instead.
No honest node can produce a mismatch. All five token-mint paths —
authenticate, signature_authenticate, api_key_authenticate, device-login
and the local auth stub — bind the caller's own key with possession proven, and
a node writes its keypair and token together in one UserSession record.
- Stop logging the node token on authentication
signature_authenticate logged the freshly issued node PASETO at INFO. The
default filter floor is info and SqliteLogger persists everything above DEBUG,
so the live bearer credential was written to the node's unencrypted log DB and
readable via fetch_logs / exportable via export_logs — the cheapest way to
harvest a node token. Log its length instead, matching the handshake path.
BIT-691
- Bill a disputed session on the customer's own claim
A disputed session billed the customer 2 * max(customer, node), so a node's
byte count set the customer's bill. Byte counts are unsigned self-claims —
BandwidthReport signs the session envelope, never the counts — and the two
peers are not equally accountable for what they assert: a customer is a KYC'd
account with a name, an email and a payment method, a node is a pseudonymous
keypair that costs nothing to discard and re-mint. A node claiming 10 GiB on a
100-byte session billed that customer ~20 GiB, with no defence available to
them: reporting truthfully does not lower a max, and reporting higher raises
it.
Bill 2 * customer_view instead. BIT-689 had already capped node credit at
min(customer, node), so both legs now rest on the customer's own figure and
the node's number moves no money in either direction. An orphan follows the
same rule: a lone customer report is a self-incriminating admission and
settles, a lone node report bills nothing.
What the node claims beyond the customer's figure is not adjudicated here —
two opposed self-reports cannot establish which is true. It accumulates into
unbilled_node_claim and is metered, being simultaneously the platform's open
liability if BIT-700's deferred settlement finds for the node and the size of
the customer-underreporting hole this opens.
BIT-689
- Move Postgres access onto sqlx, matching the fleet
BIT-692 replaced a literal NoVerifier with real certificate verification
against a mounted CloudNativePG CA, but the k8s wiring that landed with it
cannot work: the CronJob mounted a Secret named postgres-cluster-ca, which
does not exist (the real CA Secret is bitping-db-ca), in the production
namespace, when the CA lives in databases — Secrets don't cross namespaces.
Every run of the CronJob would have failed to mount and died before ever
connecting.
More fundamentally, no other service in this monorepo verifies its Postgres
TLS certificate, or hand-rolls its own rustls/tokio-postgres stack to do so.
Every other service reaches Postgres through sqlx with PgSslMode::Prefer —
TLS negotiated, neither party authenticated — and all Postgres traffic here
stays inside the cluster network, so building bespoke certificate-verification
plumbing for this one app defended against a threat nothing else here
defends against either, at the cost of infrastructure only this app needed.
Move onto sqlx entirely: a new db.rs mirrors bitping-api's services/db.rs
(PgConnectOptions + PgSslMode::Prefer, connect_lazy so the process starts
without a reachable database), SupportedCurrency derives sqlx::FromRow instead
of a hand-written tokio_postgres::Row conversion, and the fetch/insert paths
move onto sqlx::query_as/query. Drops tokio-postgres, tokio-postgres-rustls,
deadpool-postgres, rustls, rustls-native-certs, and rcgen — along with
ca_cert_path, CaCertLoad, the cnpg-ca volume mount, and the now-unbuildable k8s
wiring they existed to support.
BIT-692
-
Stamp NO_NODE_MARKER on a no-matching-node gRPC miss
-
Classify a no-matching-node miss by gRPC code, not a message substring
📚 Documentation
- Drop references to the decommissioned Bun API
The Bun service is gone. The response shapes, key orders and error envelopes
it defined are now simply the published contract that customers and the
generated SDK depend on, so the constraints stay and only the archaeology
goes. Post-cutover CLEANUP markers are removed; every test they sat on is
kept, since those tests now pin the public contract.
- Describe the corroborated-minimum node credit
The clearing rule moved from (c+n)/4 to min(c,n) per direction, but the
protocol crate's compare() doc and the game-theory README still spell out the
old formula — including the payoff matrix and the equilibrium argument that
reference it. Restate them against the rule the hub actually applies; the
conclusions are unchanged and the node's incentive to over-report is now
strictly worse than before.
🔧 Build System
- Apply pending migrations in the sqlx-cache regen
bitping-db's db/schema.sql is a snapshot of the production schema, so a
migration that has not reached production yet is absent from it and its tables
could never be type-checked by cargo sqlx prepare. Apply the migrations the
snapshot doesn't carry on top of it, so queries against a new table can be
prepared while its migration is still in review.
The snapshot stays as the baseline instead of being replaced by a full replay
because the pre-schema.sql migrations reference timescaledb, which is no
longer installed, and are not replayable from an empty database.
🚜 Refactor
- Shared logic becomes injected services
The MCP tools and the HTTP routes were about to grow two copies of the same
job dispatch, node queries and public statistics. Both now go through
JobDispatchService / NodesService / PublicStatisticsService, which own their
dependencies and are reachable from a handler as State via FromRef rather
than by naming the whole AppState.
routes/jobs/customer.rs drops to thin adapters, and the node aggregate
queries move out of routes/nodes/mod.rs into a purpose-named
routes/nodes/aggregates.rs rather than growing an already-large module.
- Node queries move into the service layer
services/nodes.rs reached back into routes/ for its SQL, so the two layers
depended on each other. The queries and their row types now live beside the
service; routes keeps the handlers, the OpenAPI annotations and the
serializers that pin the published encoding — which move to the crate root,
since both layers encode against them.
- Take auth via extractor where auth already ran first
Replaces auth::authenticate(&state, &headers) with the Auth extractor on
every handler that already called it as its first statement, and drops the
now-unused headers: HeaderMap argument where nothing else read it.
Handlers that validate a JSON body keep the plain fn: axum runs every
FromRequestParts ahead of the body-consuming extractor, so an auth
extractor there would invert the published 400-before-401 ordering. The
handlers that validate only query or path params can express that ordering,
so they get a purpose-named validation extractor declared ahead of Auth —
CountryCode on the country breakdown, ReportId on scheduled-report
delete, PayoutQuery/UnpaidQuery on the payout listings.
currency stays out of PayoutQuery because only /unpaid reads it; folding
it in would newly reject a bad value on /paid and /me, which ignore it today.
/user/me is left alone. It already resolves a JWT or an API key, but with
its own published 401 messages, and moving it onto AnyAuth would change
them.
- Make auth first and parameterise it by credential
Auth is now the FIRST thing every route runs. An unauthenticated caller must
never reach validation, or they can probe it without ever identifying
themselves. This reverses the ordering the previous commits preserved, so on
routes that validate a body or querystring a missing credential now answers
401 where it used to answer 400.
Making auth unconditionally first also removes the constraint that shaped
those commits: nothing has to run after a body extractor any more, so all ~50
handlers take the same extractor, including the 13 that were stuck on the
plain fn.
The credential is a type parameter. Auth<Jwt>, Auth<ApiKey>,
Auth<Or<Jwt, ApiKey>> — Or is a type-level combinator whose principal is
the Either enum, so a handler can match on which credential answered, and a
third credential composes as Or<Jwt, Or<ApiKey, ..>> with no new machinery.
AdminAuth<S> is the superset: same strategy, then the admin check.
Each strategy resolves its own principal rather than leaking Claims, which
is JWT-specific. Auth<ApiKey> yields a User from the key lookup;
Auth<Jwt> yields a JwtUser, which is a User plus the session a key
cannot prove. AsUser<S> collapses a richer principal for a route that never
touches the extras, so Auth<AnyUser> hands /user/me and /user/nodes a
plain User.
Ids are Uuid on the principal, parsed once in the extractor. Every
Uuid::parse_str(&claims.id) in a handler is gone, jti with them.
Parsing and validation leave the handlers entirely: PathUuid for uuid path
segments, UnpaidListing/PaidListing/MyListing yielding final limit and
page with each route's own defaults applied, and the jobs surface loses its
Caller enum now that REST arrives pre-authenticated like MCP always did.
Serde defaults cannot replace the listing extractors: ?limit= has to coerce
to 0 the way AJV's Number("") did, and f64 rejects an empty string.
The four tests pinning the old 400-before-401 ordering now pin the new one —
an anonymous caller gets 401 whatever else is wrong with the request.
- Name the freshness-window guard for the window it checks
is_later_than_5_min_from_now compared against fifteen minutes, not five, so
the one line telling a reader how far ahead of the verifier's clock a signer may
be said the wrong number.
Both bounds are now named constants — MAX_AGE_SECS and MAX_CLOCK_LEAD_SECS,
both still 15 minutes — and the two locals say which side of the window they
test. The pair of Utc::now() calls collapses into one, so both bounds derive
from a single instant rather than two readings a tick apart.
No change to either window value. The tests keep their literal 15-minute
offsets rather than referencing the new constants, so widening or narrowing a
window still has to be an explicit, reviewed test change.
- Drop the unwired credibility() primitive
credibility() (trust_weight * confidence) had exactly one caller
outside its own tests, and that caller was itself test-only
(invariants::assert_no_poison_escapes). Node selection builds on
trust_weight directly (reputation::weight::node_selection_weight),
not on credibility, so nothing in production ever called it — a
cargo check on a non-test build already flagged it dead.
Found during review of !122. Removed the function, its three
dedicated scenarios, and the now-also-unused sweep::finite_standings
helper; corrected the module doc's claim that it factors into node
selection.
🧪 Testing
- Pin the signed wire encodings
The 2026-06-11 outage was an accidental schema change to a signed wire type.
Auth<T>::verify re-encodes the decoded item with postbag::to_full_vec and
checks ed25519 over its blake3 hash, so a field added on one side of a rolling
deploy makes the two sides hash different bytes and every envelope between them
fails. Byte-exact pins turn that into a red build; round-trip tests cannot,
because encoder and decoder move together.
Pins added for the rest of the signed wire family:
Auth<T>itself — the envelope field set, the signing preimage hash
(timestamp_le || token || postbag(item)), and the base58-check signature
encoding. A change here breaks every protocol at once.- Handshake —
HandshakeRequest::Node(reaching through the proto-generated
SystemInfointoCapabilitiesandCapability, so amodels.protofield
addition lands as a test failure) and all sixHandshakeResponsevariants. - Query —
FindNodeempty and populated,FindNodes,ResolvePeers, and the
threeQueryResponsevariants. The populated cases lockRequirements,
ExclusionsandCapabilitiesfield sets, not just the discriminants. HubLoadReport.BandwidthReportandAuth<BandwidthReport>, plus a guard that
AuthedBandwidthReport's#[serde(transparent)]really costs no bytes.- The session-phase messages (
SessionInit/Ack/Transfer,TargetAddr,
the phase enums), which were round-trip-only.
wire_pin/ holds the shared harness so the two bitping-swarm pin files don't
duplicate it. The bandwidth crate's wire_snapshots.rs claimed to pin bytes
while only round-tripping; its header now says what it actually covers and
points at the new wire_pins.rs.
-
Cover the auth reject paths
-
Fuzz targets for the wire/auth decode boundary
Everything a peer sends is decoded before any signature check runs, so the
decoders are reachable by anyone who can open a stream or call the auth RPC.
Two fuzz/ crates, on the two libraries that own that boundary:
bitping-swarm — wire_decode (postbag decode of every bare wire type, plus the
re-encode fixed point Auth::verify depends on), auth_decode (the same one
layer out, on the Auth<T> envelope), auth_verify (verify end to end), and
claimed_pubkey (verify_signed_by_claimed_pubkey, which routes through the
base58 decoder).
authenticator-service — paseto_validate (all four validate entry points on
arbitrary strings) and footer_unseal (SealedFooter::try_from(Footer) and
footer_keyring::unseal, the blob carrying user_id).
Most targets carry an oracle rather than only checking for a crash: the harness
holds no private key, so anything that verifies is a forgery, and a decode that
does not re-encode to itself is the 2026-06-11 failure. footer_unseal is
deliberately panic-freedom only — its corpus contains a genuinely sealed blob to
reach the success path, so a rejection oracle there would be unsound.
Each target ships a seeds/ corpus of valid encodings plus the short and
malformed shapes worth starting from; byte mutation alone takes far too long to
synthesise a frame that decodes at all.
Kept out of the normal build: each crate declares an empty [workspace] and both
are listed under [workspace] exclude, so cargo metadata reports the same 24
members as before. cargo-fuzz needs nightly and neither it nor a nightly
toolchain is in the nix devshell, so the targets are committed unbuilt here and
not wired into CI — that needs a time budget, a corpus store and a triage owner.
docs/fuzzing.md documents the requirement and the just fuzz* recipes.
The target bodies were verified against the real APIs by compiling them as
temporary integration tests and driving 320k mutated inputs through them. No
crash on current master. Reverting the base58 fix from the previous commit makes
auth_verify reproduce that panic from wire bytes, which is the check that the
oracle has teeth.
verify_api_key_signature and hash_api_key have no target: they are
pub(super) inside a private module and reaching them would mean widening
authenticator-service's public API purely for a test harness. Noted in
docs/fuzzing.md; they have direct unit coverage instead.
- Prove the Postgres trust store refuses rogue certs
The trust-store tests covered PEM parsing but stopped short of the property
the fix exists for: that a certificate the configured CA did not issue is
refused. Drive the real rustls verifier built from build_root_store with
certificates minted at test time, and pin that a misconfigured CA path is
fatal rather than degrading to a permissive store.
root_store_adds_the_configured_ca asserted only that the store was
non-empty, which the system trust store satisfies on its own; look the
anchor up by subject instead and confirm it is absent when unconfigured.
- Pin the loopback bind for the SOCKS5 listener
The loopback default only lived in the serde attribute and the YAML comments,
so nothing failed if either drifted. Cover the default through the real
deserializer, parse the two shipped config files to confirm Config.yaml stays
on loopback while Config.docker.yaml opts into 0.0.0.0 in writing, and pin the
compose port publish that makes that opt-in safe.
- Pin multi-entry encodings for the signed collections
The existing pins were all single-entry, which is exactly why the ordering bug
was invisible: a one-element map encodes the same under any container. Adds
byte-exact pins built from out-of-sort-order insertions, plus sign → wire →
verify round-trips over 200 iterations for JobResponse::Results,
HubForwardResponse::FoundNodes, and the nested JobResult.
Each round-trip crosses a real encode/decode boundary — asserting that one
instance encodes stably would pass even with a HashMap, since the
randomisation is per-instance and the receiver's instance is the one that
differs.
Also pins the migration-safety property directly: single-entry and empty maps
must encode byte-identically under HashMap and BTreeMap.
- Prove a forged peer payload is refused on every verified path
One test per hole closed, each driving the real handler with a real
PASETO validator rather than a hand-built envelope: an authenticated peer
presents a valid token but signs the payload with a key that token does
not bind, or edits the payload after signing.
- bandwidth report: the forged and the tampered report never reach the
tracker and move no money; the honest one still bills. - FindNodes / ResolvePeers fan-out: a forged
FoundNodescontributes no
peers and resolves no routes; a correctly signed one merges. - handshake dedup: a forged
HasPeerResult { present: true }does not
deny the node, which is still registered.
Every case is paired with a negative control, because a test that only
asserts rejection passes just as happily against a handler that rejects
everything.
- Prove a harvested node token cannot claim another node's account
Drives the real HandshakeActor with a real PASETO validator. The attacker
signs honestly with its own key and declares its own key — only the token is
stolen, which is exactly the shape available to anyone who has proxied through
the victim node and read its token off the session ack.
Asserts Reauthenticate rather than acceptance, and that no
NewNodeConnection is emitted, so the victim's user_id never reaches the
pool. The negative control asserts the honest pairing still Acks, still
registers, and registers the key the token bound — without it the forgery test
would pass just as happily against a handshake that refused everything.
- Drop the wire-snapshot and map-determinism tests
Owner review on !113: BTreeMap determinism needs no test, and the wire
snapshots go with it.
- Drop the bind-host config tests
Owner review on !113.
-
Cover the claimed-pubkey length precondition
-
Pin the hard reject for an unbound handshake token
token_harvested_from_another_node_cannot_claim_its_account now asserts the
forgery is refused outright rather than invited back with Reauthenticate.
correctly_bound_token_still_acks_and_registers stays the negative control:
without it the forgery assertion would pass against a handshake that refused
everything.
- Build the authed report through its constructor
AuthedBandwidthReport's field became private in this MR so an inbound
report cannot be read without a signature check; the wire pin still built
the newtype positionally and stopped compiling. It now goes through
::new, and asserts transparency across both newtype layers.
- Spec the clearing rules as executable per-party Gherkin
The economics decide what customers pay and what nodes earn, so the rules have
to be auditable by someone who does not read Rust. clearing.feature states
every rule as a scenario grouped by whose money is at stake — what a customer
can be charged, what a node can earn, and what the platform retains and
records — and cucumber executes it, so the spec cannot drift from the
behaviour.
Replaces the hand-written test module, which rebuilt reports and outcomes at
every call site; clearing.rs drops from 598 to 207 lines. A new case now costs
five lines of Gherkin and no Rust.
Adds the cases the rewritten bill needed: an inflated node claim leaves the
customer's bill untouched, a customer admitting more than the node saw pays its
own admission, and the excess is recorded rather than billed. Two assert by
construction rather than by example — across a sweep of every claim a node
could make, none ever earns more than agreeing would, and the platform always
keeps a spread.
The hub is a binary crate, so cucumber runs from a unit test; an integration
target could not reach clear_session. It is given an explicit empty CLI because
it would otherwise parse the harness's own argv and reject cargo test <filter>, and run rather than run_and_exit keeps a failure inside the test
instead of exiting the whole binary.
- Pin the real no-node error to the real 404 end to end
The 404 tests on both sides were mock-mismatched: bitping-api hand-fed a
synthetic status into render_status and injected a pre-classified
DispatchError::NodeNotFound into its mock dispatcher, so nothing proved
the dispatcher's real error reaches the customer as a 404 — the exact gap
that let the 404 branch sit dead behind an internal status.
Drive it from the production code instead. PublicJobDispatcher::send now
runs against a real hub with an empty pool, and against a pool whose only
node fails the caller's country requirement; both have to come out of the
real ConnectionPool miss and the real GRPCJobErrors → tonic::Status
conversion as failed_precondition carrying the no-node marker. A two-hub
test wins the fanout Find phase against a node its peer hub can no longer
reach, so the Dispatch phase fails and must stay internal — the split
those two classifications must never collapse across.
bitping-api classifies and renders those same statuses through the real
render_status → dispatch_error_to_api → response bytes, and its handler
test scripts the mock with a status rather than a pre-classified error. The
marker moves to protocols::public::v1::NO_NODE_MARKER so the producing
and consuming services share one definition instead of two literals.
-
Spec the reputation ledger as executable per-party Gherkin
-
Port the test suites to executable Gherkin/cucumber specs
Converts every Rust unit/integration test file added or touched by the
BIT-689..697 + BIT-694 stack into .feature + step-definition pairs,
following the pattern established for bandwidth-report clearing and
the reputation ledger: scenarios grouped by whose money/standing is at
stake where the domain has clear parties (customer/node/platform),
topical grouping otherwise, sweep-style scenarios for grid-checked
properties rather than one scenario per grid point.
Converted, by crate:
- p2p-job-dispatcher/hub: bandwidth_report (match handling, envelope
auth), connection_pool (peer info, reputation-weighted selection),
economics (credibility, verdict), reputation store (store, subject,
flush, loader, db), integration_tests (job billing, job flow, load
shed, multi-hub), query federated-auth, peer_filter_logic,
earnings_publisher job-usage accounting. - bitping-swarm: auth verification + claimed-pubkey preconditions.
- authenticator-service: API-key login (signature verification, key
hashing; the 5 scenarios needing a real Postgres stay ignore()d
exactly as their Rust predecessors were), PASETO validation. - bitping-api: the Auth credential-extractor precedence/ordering
suite, job-dispatch error-mapping suite.
Left as plain Rust, deliberately:
- Wire-format byte pins (bitping-swarm's wire_snapshots.rs and
wire_snapshots_control_plane.rs, p2p-bandwidth-protocol's
wire_pins.rs and wire_snapshots.rs). Tried converting these first;
reverted. Each pin is one fixture in, one hex string out -- there is
no Given/When/Then narrative to gain, and wrapping a byte-exact
assert_eq! in step-matching boilerplate roughly doubled every one of
these files' line counts for zero readability benefit. A comment in
each file records this reasoning so it doesn't get re-litigated. - connection_pool/performance_tests.rs (timing/concurrency/statistical
-distribution assertions with no narrative shape). - economics/settlement.rs's own tiny arithmetic test module (three
tests, matches the precedent it was reviewed against). - tests/wire_pin/mod.rs (a shared fixture module with no test
functions of its own).
Fixed along the way: a real hang in two flush-actor scenarios (a
locally cloned mpsc::Sender outlived the loop that was supposed to
close the channel, so the flusher never observed closure and the test
never returned -- a test-harness bug introduced during this port,
caught by running it to completion rather than trusting a partial log).
- Replace the Gherkin/cucumber suites with plain rstest/proptest specs
Removes the cucumber crate and converts every one of the 27 *_scenarios.rs
step-definition modules into plain Rust tests, following the conversion
patterns settled on in review:
- Flat input->outcome rows become rstest #[case] tables; the
#[case::prose_name(...)] label carries the human-readable scenario
sentence the Gherkin Scenario: line used to. - Continuous/statistical invariants become proptest blocks, delegating to
the named asserts in economics/invariants.rs (assert_penalty_is_attenuation
etc.) so a failure names the broken rule, not the grid point. - Stateful flows (match handling, the reputation ledger, job billing)
collapse onto an ops-slice + single-driver pattern fed through the real
add_report / run_report_cleaner paths. - Libp2p/DB harness suites become thin #[tokio::test] fns over one shared
builder per module. - All 27 .feature files deleted; the cucumber dev-dep is gone from all
four manifests and its whole dependency tree is pruned from Cargo.lock.
Every scenario and every constant/tolerance is preserved verbatim; the
893-test green run and the 3/3 statistical-suite stability pass were
verified before this commit.
Fixed along the way: a flake in the flush-actor failing-sink scenario (the
shared drain() helper asserted one write per subject, but the best-effort
retry path legitimately writes more than once). The single-flush tests
keep the strict drain(); the failing-sink test now collects distinct
subjects instead - exactly what that scenario claims.