Skip to content

v0.3.0 — a user-agent is a claim; this release adds the proof

Choose a tag to compare

@Vitrus-Dev Vitrus-Dev released this 20 Sep 14:28
· 10 commits to main since this release

A user-agent is a sentence the client writes about itself. This release adds the two things that can contradict it — and is careful about what neither of them proves.

Verified agent identity

A signed request — Web Bot Auth, over RFC 9421 HTTP Message Signatures — is checked against the Ed25519 key its operator publishes, so the signer can be named. Ed25519 through Bun's WebCrypto and RFC 7638 JWK thumbprints: no dependency added, and gate:no-deps still passes.

Verification only ever promotes a label. A missing signature never makes a visitor suspicious — almost nothing signs yet, and "unsigned means fake" would just be a second unprovable claim. A key-directory miss never blocks ingest either; it returns a reason and refreshes in the background.

Agent sessions as their own class

An agentic browser (ChatGPT Atlas, OpenAI Operator) is a real Chrome session driven by an agent. It runs the tracker, clicks, fills forms and sometimes converts — and it sends an ordinary Chrome user-agent, so no bot table can name it. What it sends is Signature-Agent: "https://chatgpt.com".

Tool Behaviour What it costs
Umami, Plausible Discards it An agent that completes a checkout is revenue, and the revenue disappears
Rybbit Blocks it Same, plus you cannot tell whether agents can use your site at all
GA4 Counts it as a person Funnel, bounce rate and intent data mixed with something that is not a customer
Vitrus Its own class Excluded from visitor numbers, reported separately, with its own funnel

This count is a floor, never a total. An agent that browses without signing is indistinguishable from a person, and we do not guess. It rises only as more operators start signing.

Automation signals — the second detection layer

The loudest complaint in this category is inflated numbers: self-hosters report 200 real visitors showing up as 5,000. Plausible's cloud runs user-agent filtering plus datacentre IP ranges plus behavioural analysis; its Community Edition ships the user-agent filter alone. Ours was one layer too.

Five weighted header rules — no-accept-language (3), no-fetch-metadata (3), chromium-without-client-hints (3), platform-mismatch (5), no-accept (2) — with a threshold of 5, so no single rule can accuse anyone: one missing header is normal somewhere on the internet. Applied only to clients claiming to be a browser; something that honestly says curl/8.4 is already named by the user-agent table. All of it in the Apache-2.0 core.

Recorded, not enforced. The rules that fired are written onto the event (bot_signals, bot_score) and change no number. Suspected traffic is still counted everywhere until you ask otherwise. The exclusion is not remembered between visits, a banner stays on screen while it is on, and the evidence panel shows the predicate doing it:

SELECT COUNT(DISTINCT visitor_id) AS value
  FROM events
 WHERE site_id = ? AND ts >= ? AND ts < ?
   AND bot_kind = '' AND agent_trust = 'human'
   AND bot_score < 5

A silent filter is indistinguishable, from your side, from a bug that loses traffic. "Rather undercount than guess" cuts both ways: we will not guess that someone is a robot either.

Not a fingerprint. No canvas, font enumeration, WebGL or plugin list — the usual meaning of a "client signals" layer, and all of it would undercut the reason to run this instead of Google Analytics. Every rule reads a header the client sent anyway, and a test fails the build if a fingerprinting surface ever appears in that module.

Known blind spots, since a detection page that lists none is selling something: a headless browser with stealth patches sends plausible headers and scores zero, and so does a residential-proxy botnet running real Chrome. The rule table is versioned and will grow; it will not reach certainty.

Also in this release

  • Filters compiled into the SQL on the server rather than applied to rows in the browser, so the query you can read is the query that produced the number.
  • Country centroids in the core, for country-level mapping without a GeoIP database.
  • The self-hosted dashboard gained a catch-all for scalar metrics. It already had one for rows, so a scalar the engine computed but the card list did not name was simply invisible — the open-core claim held for only half the bundle.

One correction

The bot-detection documentation previously said header heuristics were "not a trade available to us". Adding them made that false, so it is retracted in place rather than quietly rewritten. What stays true is the reason the two are kept apart: a signed agent is a fact, a suspicious one is an opinion, and summing them produces a number that means neither.


315 tests, two golden-set evals at 100%, four build gates. Upgrading is a restart: the new columns are added by an additive migration and existing rows default to "nothing was observed", which is the honest reading.

Full changelog: v0.2.0...v0.3.0