Skip to content

Pulse v0.2.0

Choose a tag to compare

@github-actions github-actions released this 29 Sep 12:37
· 118 commits to main since this release
74005f6

The first stable release since 0.1.0. Wake up, admin: your server has a pulse, and now it can tell you which mod is spending it.

This tag folds in everything published across the five 0.2.0 prereleases (indev.1 through indev.5), plus a round of fixes that landed after indev.5. If you already run indev.5, read "Fixed since indev.5" below. If you are upgrading from 0.1.0, read the whole thing: there is one security note and three upgrade hazards that need your attention, a Bind: 0.0.0.0 setting that finally takes effect, a breaking metric rename and, if you push OTLP, a change to the job and instance labels your backend assigns each series.

Security note: Docker commands in earlier READMEs

If you ever ran the docker run commands from an earlier version of contrib/grafana/README.md or contrib/alerts/README.md, read this first. Those commands started Grafana and Prometheus with host networking and no bind to loopback, and Grafana with anonymous admin access on top. On Linux, that Grafana gives administrator access to anyone who can reach port 3000 on the machine, and Prometheus serves your server's metrics on port 9090 the same way. Stop both containers (docker stop pulse-graf pulse-prom), then start them again from an up-to-date copy of the kit: docker compose up -d in contrib/grafana, or, if you load the alert rules, the current docker run commands in contrib/grafana/README.md and contrib/alerts/README.md. All of them bind to 127.0.0.1 only. To view a remote server's dashboard, use an SSH tunnel; docs/getting-started.md shows how.

Upgrading from 0.1.0

Replace pulse_0.1.0.zip in Mods/ with pulse_0.2.0.zip, do the same for pulseotlp if you push OTLP, and restart once. Nothing needs hand-editing: each mod checks its config file at startup and writes back any key it is missing, with that key's default, so pulse.json picks up the new Attribution block and pulse-otlp.json picks up ServiceName on their own. Values you already set are left exactly as they are. A config key neither mod recognises is named in a warning instead of disappearing quietly, since the rewrite drops it. That rewrite only happens when a key is missing, and it writes a fresh copy of the file from the loaded values, so a comment you hand-added to either file will not survive one.

If your pulse.json still sets Bind to 0.0.0.0 from an earlier attempt, check it before you restart: that value failed to bind under 0.1.0 (Pulse logged an error at every start and served nothing), so an admin who set it and moved on may never have noticed. The plain-socket endpoint (see What's new) binds it, so this restart is the one where 0.0.0.0 starts working, serving /metrics on every interface with no authentication, exactly what 0.1.0's own README warned against.

Every 0.2.0 start also logs the engine's own warning, "Over 400ms tick. Skipping N physics ticks.", once at startup, since Pulse primes the engine's frame profiler for the first tick after every start (attribution on or off) and the engine only prints that line while its profiler is on; it adds 1 to pulse_log_entries_total{level="warning"} and is otherwise harmless, since it is not one of the kinds pulse_engine_warnings_total counts, so no bundled alert fires over it. With attribution on, the same line also appears during a profiled burst whenever physics falls behind.

If you push OTLP, this upgrade also changes the service.name your backend sees, and with it the job label Prometheus's OTLP receiver, Mimir and Grafana Cloud derive from it, plus an instance label 0.1.0's series never carried. 0.1.0 never set service.name itself, so the OpenTelemetry SDK exported its own fallback, unknown_service: followed by the process name (dotnet when launched as dotnet VintagestoryServer.dll, VintagestoryServer through the native apphost), unless you had already set OTEL_SERVICE_NAME or OTEL_RESOURCE_ATTRIBUTES=service.name=...: the SDK's own default resource already reads both, with no ServiceName key involved and no instance label.

Unless OTEL_SERVICE_NAME is set, 0.2.0 calls AddService, and two things follow. Left at its default, ServiceName changes every existing series' job to vintagestory, so a dashboard variable or alert keyed on the old value breaks quietly rather than erroring. Whatever ServiceName holds, every series also gains an instance label, from a service.instance.id that AddService regenerates at random on every restart.

Setting ServiceName to your old value brings the old job back but not the old identity: you still get the new, restart-churning instance label. To keep 0.1.0's job and instance labels exactly, set OTEL_SERVICE_NAME to your old value instead: the mod then skips AddService entirely, so neither job nor instance changes. If you already had OTEL_SERVICE_NAME set under 0.1.0, neither job nor instance changes either, though target_info's telemetry_sdk_version moves from 1.18.0 to 1.19.1 and the resource gains a schema URL.

A service.name or service.instance.id set through OTEL_RESOURCE_ATTRIBUTES is silently overridden by AddService unless OTEL_SERVICE_NAME is also set. Set OTEL_SERVICE_NAME to move your service.name there; the service.instance.id from OTEL_RESOURCE_ATTRIBUTES then survives, since that also skips AddService.

For a stable instance label across restarts, set both: OTEL_SERVICE_NAME=<name> together with OTEL_RESOURCE_ATTRIBUTES=service.instance.id=<id>. OTEL_RESOURCE_ATTRIBUTES alone is not enough, since AddService's own randomly generated id silently overrides it on every start.

If you ever roll back to 0.1.0, it leaves those config files exactly as they are: it does not know Attribution or ServiceName and does not strip them, it just ignores them, so nothing on disk is lost, but any attribution tuning goes inert until you upgrade again, and so does ServiceName: 0.1.0 ignores that key and goes back to the SDK's own unknown_service: fallback job, unless OTEL_SERVICE_NAME is set. 0.1.0 honours that variable through the same SDK default resource 0.2.0 does, so a name set that way carries through the rollback unchanged; a name set only through the ServiceName config key does not.

What's new

  • Per-mod tick attribution. Turn it on and see which mod is actually spending the tick, on a live graph rather than a one-off report: pulse_mod_tick_share{modid}, pulse_mod_tick_seconds_total{modid}, and two counters behind them. Off by default; see the cost note below before you turn it on.
  • /pulse, a new server command (behind controlserver). /pulse attribution on|off|status switches the duty cycle on a running server with no restart, without writing pulse.json, so a switch does not survive one, and /pulse reload re-reads pulse.json and applies the Attribution block live, naming any other key that still needs a restart to take effect.
  • The metrics endpoint is now a plain socket, not HttpListener. Binding loopback on Windows no longer needs administrator rights or a netsh reservation, a scrape of http://localhost:9464/metrics on Linux no longer gets a 404 because its Host header does not match Bind, several scrapes are served at once instead of queued, and Bind: 0.0.0.0 now binds instead of failing with a logged error.
  • Config files upgrade and tolerate damage. Covered above for the happy path; a pulse.json or pulse-otlp.json that exists but will not parse (a doubled comma, a missing quote) no longer stops the mod from starting. Each mod logs the file's path and the parser's own message and leaves the file untouched; Pulse runs that session on its built-in defaults, and Pulse OTLP leaves export off for that session rather than guess an endpoint.
  • OTLP export failures no longer pass silently. A rejected push, a refused or unreachable collector, a timeout: each kind now logs one line, repeated at most every 10 minutes, naming what failed and the backend's own (redacted) response. A matching line reports the first successful export after a failure, or the first export after each server start when nothing has failed yet. Six families that could stay invisible on a quiet server until their first real event (engine warnings, player deaths, suspends, suspend seconds, worldgen columns, log entries) now reach OTLP seeded at zero from the first export instead of waiting. A new ServiceName config key (default vintagestory) sets the service.name resource attribute; set it per server so a backend collecting from more than one can tell them apart, since the default is the same on every install. OTEL_SERVICE_NAME still overrides it if you set that instead.
  • contrib/alerts, a new Prometheus alerting rules pack, eleven rules covering tick rate, tick saturation, sustained tick overruns, engine warnings, log errors, endpoint availability and a stuck worldgen queue, one of them PulseModHoggingTick, which only fires when a single mod holds more than half the profiled tick for 10 minutes while the server is also over 80% of its tick budget. The bundled Grafana dashboard gains a matching attribution row.
  • docs/getting-started.md, a walkthrough for a server owner who has never used Prometheus or Grafana, routed by how the server is hosted, ending at the shared dashboard either way.

Breaking change: nine runtime series renamed

Prometheus, Mimir and Grafana Cloud derive a name from an OTLP instrument by looking at its unit. Nine dotnet_* families on /metrics used to skip that step, only mapping dots to underscores and appending _total, so they were out of step with every other OTLP-derived name on the same dashboard. From 0.2.0 they match:

Old name New name
dotnet_process_memory_working_set dotnet_process_memory_working_set_bytes
dotnet_gc_heap_total_allocated_total dotnet_gc_heap_allocated_bytes_total
dotnet_gc_last_collection_memory_committed_size dotnet_gc_last_collection_memory_committed_size_bytes
dotnet_gc_last_collection_heap_size dotnet_gc_last_collection_heap_size_bytes
dotnet_gc_last_collection_heap_fragmentation_size dotnet_gc_last_collection_heap_fragmentation_size_bytes
dotnet_gc_pause_time_total dotnet_gc_pause_time_seconds_total
dotnet_jit_compiled_il_size_total dotnet_jit_compiled_il_size_bytes_total
dotnet_jit_compilation_time_total dotnet_jit_compilation_time_seconds_total
dotnet_process_cpu_time_total dotnet_process_cpu_time_seconds_total

pulse_* families are unaffected. If a panel or alert of yours queries one of the nine, update it; the bundled dashboard queries both the old and new name in this release, so it stays populated whether the Pulse it points at has upgraded yet or not, and you can drop the fallback side once every server you watch is on 0.2 or later.

A smaller, related fix: pulse_network_packets_per_second and pulse_network_bytes_per_second had picked up a doubled _per_second suffix over OTLP ever since they shipped in 0.1.0. That is corrected too. /metrics never carried the doubled name, so this only matters if you query a Grafana Cloud (or other OTLP backend) series under the old, doubled name; historical data stays under it, new data lands under the corrected one.

Fixed since indev.5

  • Pulse no longer switches the engine's frame profiler off on every tick. indev.1 never touched the profiler at all; indev.2 and indev.3 only wrote it every tick while attribution was turned on; indev.4 and indev.5 wrote it every tick even with attribution off, which broke /debug logticks (its per-system and per-listener lines went missing, off-thread reports stopped) and overrode any other mod that turns the profiler on.
  • Pulse OTLP no longer crashes on unusual config. An Endpoint with an empty user name and a password (http://:token@collector:4318) took the whole game server down at boot, in 0.1.0 as well. A header value containing a comma, or two header names that collide once trimmed, stopped the mod with a stack trace. Each now logs one error, never the value, and leaves export off. IntervalSeconds is now capped at 86400, a day, so a value several digits too long no longer overflows.
  • A missing config file on a read-only ModConfig folder no longer stops either mod: Pulse runs the session on its defaults, Pulse OTLP with export off.
  • A /pulse command already registered by another mod is now logged once instead of aborting Pulse's startup with the frame profiler stuck on. A game update that reshapes the frame profiler now logs one warning and turns attribution off for the session, instead of crashing startup or logging an error every tick.
  • The export-failure log now also redacts the query parameter values and userinfo of your own Endpoint, so a proxy error page that echoes a signed URL should no longer put its secret in the server log; redaction has a 6 character floor and none of this is exhaustive.
  • Smaller fixes: a ChunksRefreshSeconds above about 2.1 million no longer makes the chunk read run every tick. Config keys written in another case (port for Port) are no longer reported as both added and dropped, a key duplicated under two casings is reported without printing its value, and the config upgrade's log lines and /pulse attribution off's reply now say what actually happened.
  • pulse_server_tick_seconds's histogram gains two bucket boundaries, at 0.035 and 0.04 seconds, just above the 33.3 ms tick budget. A healthy server ticking only a shade slow no longer shows a misleading p99 near 50 ms on the bundled dashboard; every existing boundary is unchanged.

Attribution costs more than the early prereleases said

If you turned attribution on under indev.2, indev.3 or indev.4, this correction is for you. The first estimate, from counting profiler marks, put the amortised cost at about 0.3% of the tick budget. Measured instead, with a scenario that spawns four thousand entities and reads real tick busy time across matched on/off windows, a profiled tick costs about 26% of the budget while a burst runs. At the shipped default (BurstTicks: 10, every 10 seconds) that blends down to about 0.9% amortised; the previous default of 30 ticks, shipped in those builds, amortised to about 2.5%. A pulse.json written by an earlier prerelease keeps whatever BurstTicks it was written with: set it to 10 and run /pulse reload to pick up the new default without a restart. Attribution is still off unless you turn it on, so none of this costs you anything until you do.

Compatibility

Vintage Story 1.22.x, dedicated server only, .NET 10, Linux and Windows.

Test bar for this tag: 484 unit tests and 39 scenarios on a real headless server, mutation scores of 94% (Pulse) and 81% (Pulse OTLP), both above their break thresholds, and a passing SonarCloud quality gate. On top of that, a final end-to-end run built the real release zips and drove them against a real dedicated server, eighteen checks in all: build and packaging; config file creation, upgrade and an unreadable file; /metrics plus promtool; the /pulse console commands; Prometheus scraping and Grafana provisioning the bundled dashboard, with each of its 50 panel queries run against both the direct scrape and the OTLP-fed Prometheus; all eleven contrib/alerts rules loading and none firing; the Host header and Bind: 0.0.0.0 cases; scrapes served while other connections are held open; OTLP metric-family parity against a real Prometheus OTLP receiver; OTLP export-failure logging against a closed port, and header redaction against a mock collector answering 401; and a clean shutdown. Every check passed. Three of them, package, 5b and 7b, also carried an informational note: the release zips' own file sizes, and one or two dashboard queries with no data yet at the moment they ran, since no .NET exception had been thrown during the run and the slow-cadence entities-by-code gauge had not yet reported over OTLP. Check 1b passed too; its promtool exit code was lint only (three _count suffix warnings on OpenTelemetry's dotnet_* runtime gauges).

Everything else, including the smaller fixes, is in the changelog.