Repository navigation
Releases: Vitrus-Dev/vitrus
Release list
v0.6.0 — a bigger MCP server
MCP
- query_stats — any metric (visitors, sessions, pageviews, events) by any of 20 dimensions or by day, with filters and exact from/to dates. Returns the rows and the SQL.
- get_realtime, get_revenue (per currency, never converted), get_journeys, list_goals / goal_report, get_tracking_snippet.
- Every tool declares
readOnlyHint; ranges acceptdaysorfrom/to. - The host can lend saved goals and the tracker origin through
ToolContext.
15 tools in total, all read-only, every answer with its evidence.
v0.5.1 — self-hosted dashboard password
Security
VITRUS_PASSWORD— set it and the self-hosted dashboard, its read API and the replay settings ask for it (HTTP Basic, any user name). The tracker and ingest stay public, because your visitors' browsers call them. Until now these were open to anyone who could reach the port;vitrus startnow says so when the variable is unset. Upgrade and set it if your server is reachable from outside.- Ingest bodies over 64 KB are refused with 413 (Bun's default limit was 128 MB).
Docs
- README rebuilt around screenshots of the hosted dashboard on a demo workspace (sample data).
0.5.0 — revenue and Umami import in the core
Revenue
Send an amount and a currency with any event:
vitrus("purchase", { revenue: 49.90, currency: "USD" })- Normalised at ingest.
revenuemust be a number ≥ 0 (a numeric string such as"49.90"is accepted),currencya three-letter ISO 4217 code in any case. - A bad amount never loses the event. An amount that is not a number, negative, over one billion or missing its currency is removed, and the reason is kept in
_revenue_rejectedso it can be counted and shown rather than silently reported as less. computeRevenuereturns, for one currency at a time: revenue, orders, average order value, revenue per visitor (over all human visitors), conversion rate (sessions with a payment ÷ all sessions), a breakdown by channel, source, campaign, event, page, country or device, a series, the previous period and the rejected amounts by reason — each as a query with its SQL and parameters.- No currency conversion. Amounts in different currencies are never added together: an exchange rate is a number from somewhere else, and a total built on one could not be checked against the data.
- Attribution. Channel, source and campaign are where the visit began — the session's first event — because a checkout page almost never has an outside referrer. Page, event, country and device are the payment's own.
- Revenue from verified AI-agent sessions is reported on its own line, not in the totals.
Import from Umami
importUmamiCsv(store, csv, options) reads an Umami export — one row per page view or event, the same grain Vitrus stores — into events.
- Idempotent. Each event keeps its Umami id; importing the same file twice adds nothing, and the summary reports what was actually written.
- No double counting. Rows at or after the first event Vitrus itself collected are skipped and counted.
- Same rules. Channels and sources are re-classified with Vitrus's referrer table. Umami's session becomes a visitor and its visit a session.
- Removable. Every row carries its batch id;
removeImportdeletes exactly one batch. - Plausible is deliberately not supported: its export contains daily totals, not visits, and turning totals into visits would mean inventing them.
Scope
Both are in the core library. The hosted dashboard at app.vitrus.dev has a Revenue page and an upload for imports; the self-hosted server does not expose routes for them yet — call the functions directly for now.
0.4.2 — ingest at /api/d, and a header for proxies
The tracker posts to /api/d
Content blockers match /api/collect as a path on any domain — first-party included. A visitor running one sent nothing at all, even to a self-hosted install on the site's own domain. The tracker now sends events to /api/d.
/api/collect is unchanged and remains the same endpoint, so existing tags, server-side integrations and anything calling it directly keep working. A page still serving a cached older v.js keeps posting there until its cache expires.
X-Vitrus-Client-IP for proxies
If you serve the tracker through a proxy on your own domain, every request reaches the server from the proxy's address — and with Cloudflare in front, cf-connecting-ip is the proxy, not the visitor. All visitors then hash to one.
The client IP is now read in this order: X-Vitrus-Client-IP → cf-connecting-ip → the first value of X-Forwarded-For → X-Real-IP. Set the first one in your proxy:
location = /vs/api/d {
proxy_pass https://your-vitrus-host/api/d;
proxy_set_header X-Vitrus-Client-IP $remote_addr;
}Like X-Forwarded-For, it can be set by the client. The IP only feeds a daily-salted visitor hash and IP exclusions; it is never stored and never used for an access decision.
Full proxy guide (nginx, Caddy, Cloudflare Workers, Next.js): https://vitrus.dev/docs/proxy
Fix
- The session-replay recorder URL was derived by trimming a fixed number of characters off the collect URL. It is now built from the endpoint base, so
data-hostand the new path both resolve correctly.
0.4.1 — client IP behind Cloudflare
Fix: one visit counted as several visitors behind Cloudflare
The client IP now comes from cf-connecting-ip when it is present, and only then from the first value of x-forwarded-for.
Behind Cloudflare and a reverse proxy that rewrites x-forwarded-for (Caddy does this by default for clients it does not trust), the first value of that header is the Cloudflare edge that connected, not the visitor, and it changes from request to request. Because the anonymous visitor id is a daily-salted hash of IP and user-agent, a single page visit — its pageview, an outbound click, and the Web Vitals beacon — could be hashed into several visitors, each with its own zero-length session.
Symptoms you may have seen: sessions roughly doubled, pages-per-session under 1, a bounce rate near 100%, and funnel steps after the first page empty.
What this does not fix: events already stored. Visitor ids are one-way hashes, so past sessions cannot be re-joined. Numbers are correct from the moment you upgrade.
If you run behind Cloudflare without Caddy, or with trusted_proxies configured, you were not affected.
0.4.0 — globe, real-time, sessions, journeys, goals and opt-in replay
Analytics that shows its work. Every number, list and chart added in this release opens the SQL that produced it — and the AI digest still drops any sentence whose number is not in that evidence.
What's new
- Globe, 3D and 2D. Sessions placed at their city on a timeline you can scrub and play, coordinate cells, and choropleths by country and by state or province. City, region and 0.1° coordinates come from your proxy's headers (Cloudflare's visitor-location transform, Vercel, CloudFront). The geometry is served from your own origin: no map provider, no token, no IP lookup.
- Real-time. A live count on every page and a Realtime view with views per minute and the latest events labelled person, crawler or agent. Server-sent events with a polling fallback.
- Sessions, users and events. A session timeline per visit, user profiles for people you
identify()(the raw id is hashed before it is stored), event property breakdowns, and errors grouped by message, file and line. - Journeys and goals. Path queries for a Sankey, goals on page wildcards, custom events or event properties, and ordered funnels.
- Explore queries. Filters take is / is not / contains / starts with / regex on 24 dimensions and any event property, compiled into the SQL on the server.
- Performance. p50 through p99 for five Web Vitals.
- Session replay, opt-in. Off by default, loaded only on pages that carry
data-replay, and every text node and input value masked unless you unmask an element on purpose. Password, card and one-time-code fields are never captured, DNT and GPC stop recording, and the server masks again on ingest. The player runs in a script-less sandboxed iframe. It is a hand-written recorder at 3.6 KB gzipped in a separater.js, with its own 5 KB CI budget, so pages without it keep the 2.7 KB tracker.
Caveats
- Cities need your proxy to send them. Without those headers the globe names the missing header instead of drawing an empty map, and country-only sessions are never pinned to a country's centre.
- Journeys, goals and user profiles join days together only for people you identify. An anonymous visitor's id rotates daily by design.
- The tracker grew by about 200 bytes gzipped (file downloads, hash routing, user traits). The 3 KB gate still holds.
- The self-hosted dashboard (
vitrus start) renders every metric and plays replays; the sessions, users, journeys, goals and globe pages are in the hosted dashboard today. Their queries are in this core.
Full changelog: https://vitrus.dev/changelog
v0.3.0 — a user-agent is a claim; this release adds the proof
A user-agent is a sentence the client writes about itself. This release adds the two things that can contradict it — and is careful about what neither of them proves.
Verified agent identity
A signed request — Web Bot Auth, over RFC 9421 HTTP Message Signatures — is checked against the Ed25519 key its operator publishes, so the signer can be named. Ed25519 through Bun's WebCrypto and RFC 7638 JWK thumbprints: no dependency added, and gate:no-deps still passes.
Verification only ever promotes a label. A missing signature never makes a visitor suspicious — almost nothing signs yet, and "unsigned means fake" would just be a second unprovable claim. A key-directory miss never blocks ingest either; it returns a reason and refreshes in the background.
Agent sessions as their own class
An agentic browser (ChatGPT Atlas, OpenAI Operator) is a real Chrome session driven by an agent. It runs the tracker, clicks, fills forms and sometimes converts — and it sends an ordinary Chrome user-agent, so no bot table can name it. What it sends is Signature-Agent: "https://chatgpt.com".
| Tool | Behaviour | What it costs |
|---|---|---|
| Umami, Plausible | Discards it | An agent that completes a checkout is revenue, and the revenue disappears |
| Rybbit | Blocks it | Same, plus you cannot tell whether agents can use your site at all |
| GA4 | Counts it as a person | Funnel, bounce rate and intent data mixed with something that is not a customer |
| Vitrus | Its own class | Excluded from visitor numbers, reported separately, with its own funnel |
This count is a floor, never a total. An agent that browses without signing is indistinguishable from a person, and we do not guess. It rises only as more operators start signing.
Automation signals — the second detection layer
The loudest complaint in this category is inflated numbers: self-hosters report 200 real visitors showing up as 5,000. Plausible's cloud runs user-agent filtering plus datacentre IP ranges plus behavioural analysis; its Community Edition ships the user-agent filter alone. Ours was one layer too.
Five weighted header rules — no-accept-language (3), no-fetch-metadata (3), chromium-without-client-hints (3), platform-mismatch (5), no-accept (2) — with a threshold of 5, so no single rule can accuse anyone: one missing header is normal somewhere on the internet. Applied only to clients claiming to be a browser; something that honestly says curl/8.4 is already named by the user-agent table. All of it in the Apache-2.0 core.
Recorded, not enforced. The rules that fired are written onto the event (bot_signals, bot_score) and change no number. Suspected traffic is still counted everywhere until you ask otherwise. The exclusion is not remembered between visits, a banner stays on screen while it is on, and the evidence panel shows the predicate doing it:
SELECT COUNT(DISTINCT visitor_id) AS value
FROM events
WHERE site_id = ? AND ts >= ? AND ts < ?
AND bot_kind = '' AND agent_trust = 'human'
AND bot_score < 5A silent filter is indistinguishable, from your side, from a bug that loses traffic. "Rather undercount than guess" cuts both ways: we will not guess that someone is a robot either.
Not a fingerprint. No canvas, font enumeration, WebGL or plugin list — the usual meaning of a "client signals" layer, and all of it would undercut the reason to run this instead of Google Analytics. Every rule reads a header the client sent anyway, and a test fails the build if a fingerprinting surface ever appears in that module.
Known blind spots, since a detection page that lists none is selling something: a headless browser with stealth patches sends plausible headers and scores zero, and so does a residential-proxy botnet running real Chrome. The rule table is versioned and will grow; it will not reach certainty.
Also in this release
- Filters compiled into the SQL on the server rather than applied to rows in the browser, so the query you can read is the query that produced the number.
- Country centroids in the core, for country-level mapping without a GeoIP database.
- The self-hosted dashboard gained a catch-all for scalar metrics. It already had one for rows, so a scalar the engine computed but the card list did not name was simply invisible — the open-core claim held for only half the bundle.
One correction
The bot-detection documentation previously said header heuristics were "not a trade available to us". Adding them made that false, so it is retracted in place rather than quietly rewritten. What stays true is the reason the two are kept apart: a signed agent is a fact, a suspicious one is an opinion, and summing them produces a number that means neither.
315 tests, two golden-set evals at 100%, four build gates. Upgrading is a restart: the new columns are added by an additive migration and existing rows default to "nothing was observed", which is the honest reading.
Full changelog: v0.2.0...v0.3.0
v0.2.0 — the dashboard shows the whole engine
The first tagged release. 0.1.0 was the initial publication of the core and was never tagged; this is the same engine with a schema fix, a dashboard that shows all of it, and a fourth build gate.
Fixed — upgrading a self-hosted database could break adding a site
sites.privacy_mode was added to the schema after the table first shipped. CREATE TABLE IF NOT EXISTS is a no-op on a table that already exists — it does not reconcile columns — so the column was present on every fresh install and on none of the upgraded ones. The gap only surfaced on write, as an opaque 500 internal_error with nothing in the UI to explain it.
If you self-host and installed before that column existed, adding a site has been failing. Upgrading to 0.2.0 fixes it on the next start: SqliteStore.init() now checks PRAGMA table_info and adds missing columns explicitly. No data is touched — the existing rows keep their values and get the default.
Four tests build a pre-column database and pin the behaviour, including a double init() so a restart cannot crash on boot.
The self-hosted dashboard now shows the whole engine
It rendered six cards and the digest. The engine had always computed far more, and the documentation said so — which made "no metric is held back from this repository" look false to anyone who actually checked. Only the screen was missing.
Now on the page, each with the evidence control it always had:
- Performance — Core Web Vitals at p75, slowest pages
- Errors — grouped by message and page, and by browser and OS
- Traffic — top, entry and exit pages, channels, referrers, UTM campaigns
- AI — AI sources, their landing pages, and which of your pages AI crawlers read
- Behaviour — custom events, field-level form abandonment
- Audience — country, device, browser, OS, language, screen
The section list is derived from the bundle at runtime, so a metric added to metrics/queries.ts appears without anyone remembering to update the page, and a metric no section claims renders under More rather than vanishing.
New gate: gate:inline-js
The dashboard's JavaScript is generated inside a TypeScript template literal. One mis-escaped quote kills the entire script while the page still renders perfectly — no error, no console trace, nothing works. Every text-matching test passes, because the text is in the HTML.
The gate parses every inline script instead of searching it. It caught a real mismatched quote while this release was being written.
README
- A dashboard preview, generated rather than screenshotted. A screenshot goes stale the first time a label moves and nobody notices for months — and a screenshot full of invented numbers presented as real would be the worst possible advertisement for a product whose argument is that numbers should carry their evidence. The word "illustration" is inside the artwork, not in a caption that a re-screenshot would crop away.
- A dated comparison against GA4, Plausible, Umami, Rybbit and Databuddy — including the rows where each of them is the better choice, and a note that they all classify AI assistants as a channel. The narrower thing we are claiming is separating the crawler that read you from the human it sent.
- A table of contents, community links, and correct test counts.
The table is deliberately short: it lists only differences that can be stated without qualification. An earlier draft claimed Umami had no MCP server; checking found their official one. A half-verified cell is exactly the kind of claim this project exists to stop making.
Upgrading
git pull && bun install && bun run build:trackerSchema migrations run on start. Nothing else to do.
198 tests, two golden-set evals and four gates, as one command: bun run gates.