Releases: StratumServer/Pulse
Release list
Pulse v0.2.1
The second stable release of the 0.2 line. It is v0.2.1-indev.1 unchanged, and it brings three changes since 0.2.0. Two of them can touch what you built around Pulse: if your server's log is shipped somewhere, or its OTLP series feed a dashboard or an alert, read the first two sections before you upgrade.
Log lines now carry the mod's id
Both mods now write through the logger the game gives each mod, so every line they write has the mod's id right after the severity. [Notification] Pulse serving metrics on http://127.0.0.1:9464/metrics becomes [Notification] [pulse] Pulse serving metrics on http://127.0.0.1:9464/metrics, and Pulse OTLP's lines carry [pulseotlp]. The messages themselves are word for word what they were. A pattern in a log shipper or an alert rule that expects Pulse's words straight after the severity needs the tag in between. pulse_log_entries_total and pulse_engine_warnings_total count exactly what they counted before.
Pulse OTLP: one instance label per server, across restarts
0.2.0 exported a service.instance.id that the OpenTelemetry SDK generated at random on every start (unless OTEL_SERVICE_NAME was set). Prometheus, Mimir and Grafana Cloud turn it into the instance label, so every restart started a new set of series, and an alert keyed on instance saw a new server each time.
The id now lives in a new ServiceInstanceId key of pulse-otlp.json. Left blank, which is what an upgraded file has, it gets a generated GUID on the first start, written into the file, and every start after that reuses it. That first start changes the label one last time. You can also put a readable id in the key yourself, survival-eu-1 say. The first start logs Pulse OTLP wrote these keys into pulse-otlp.json: ServiceInstanceId. Everything else in the file was kept as it was., and the startup line now names the id: Pulse OTLP exporting ... as service 'vintagestory', instance '<id>'.
Check three things when you upgrade:
- Each server needs an id of its own. A
pulse-otlp.jsoncopied to a second server, or a template several servers'ModConfigfolders are built from, gives all of them one id, and two servers with the sameServiceNameand the same id are one server to a backend. Clear the key in the copies, or give each server its own value. - Pulse OTLP cannot save the generated id when
ModConfigis read-only. It then logs a warning naming the id and saying the next start will export a different one. PutServiceInstanceIdin the file yourself, or setOTEL_RESOURCE_ATTRIBUTES=service.instance.id=<id>in the server's environment. AModConfigthat does not survive a restart (a container without a volume for it) loses the id the same way, with no warning. - The first start rewrites
pulse-otlp.jsonto add the key. Like every config upgrade, the rewrite drops comments in the file, and names any key Pulse does not know in a warning before dropping it.
The environment still has the last word. A service.instance.id in OTEL_RESOURCE_ATTRIBUTES wins over the key, and it now survives on its own, where 0.2.0 replaced it on every start unless OTEL_SERVICE_NAME was set too. With OTEL_SERVICE_NAME set, nothing changes: the environment decides the whole identity, as before.
Attribution: the game's own work no longer lands in unattributed
This one only matters with attribution on. Part of the game's own entity behaviours (despawn timers, name tags, creature physics and more) was filed under unattributed, because the name a behaviour marks the tick with is often not the code its class was registered under. Pulse now reads the loaded entities to learn which mod ships each behaviour, and that time goes to game, survival or engine. On a test server with 1,691 entities, unattributed went from about 2.5% of the sampled tick to under 0.1%.
The read happens outside the profiled ticks. The first burst after attribution starts reads every loaded entity: about 3 to 4 ms with 1,700 entities, 7 ms with 4,000, 23 to 40 ms with 8,000. After that a burst costs about 0.1 ms, plus the entities that loaded in between.
Documentation
The getting-started guide now covers macOS, gives the right data path for the server.sh that ships with the server (/var/vintagestory/data/Mods), and explains how to load the bundled alert rules into Prometheus.
Upgrading
Replace pulse_0.2.0.zip in Mods/ with pulse_0.2.1.zip, do the same for pulseotlp if you push OTLP, and restart. pulse.json needs no change; pulse-otlp.json gains its ServiceInstanceId on that first start, as described above. Coming from v0.2.1-indev.1, swap the zips: the code is the same.
Full changelog: v0.2.0...v0.2.1
Pulse v0.2.1-indev.1
The first prerelease of 0.2.1. It brings three changes since 0.2.0, and two of them can touch what you built around Pulse: if your server's log is shipped somewhere, or its OTLP series feed a dashboard or an alert, read the first two sections before you upgrade.
Log lines now carry the mod's id
Both mods now write through the logger the game gives each mod, so every line they write has the mod's id right after the severity. [Notification] Pulse serving metrics on http://127.0.0.1:9464/metrics becomes [Notification] [pulse] Pulse serving metrics on http://127.0.0.1:9464/metrics, and Pulse OTLP's lines carry [pulseotlp]. The messages themselves are word for word what they were. A pattern in a log shipper or an alert rule that expects Pulse's words straight after the severity needs the tag in between. pulse_log_entries_total and pulse_engine_warnings_total count exactly what they counted before.
Pulse OTLP: one instance label per server, across restarts
0.2.0 exported a service.instance.id that the OpenTelemetry SDK generated at random on every start (unless OTEL_SERVICE_NAME was set). Prometheus, Mimir and Grafana Cloud turn it into the instance label, so every restart started a new set of series, and an alert keyed on instance saw a new server each time.
The id now lives in a new ServiceInstanceId key of pulse-otlp.json. Left blank, which is what an upgraded file has, it gets a generated GUID on the first start, written into the file, and every start after that reuses it. That first start changes the label one last time. You can also put a readable id in the key yourself, survival-eu-1 say. The first start logs Pulse OTLP wrote these keys into pulse-otlp.json: ServiceInstanceId. Everything else in the file was kept as it was., and the startup line now names the id: Pulse OTLP exporting ... as service 'vintagestory', instance '<id>'.
Check three things when you upgrade:
- Each server needs an id of its own. A
pulse-otlp.jsoncopied to a second server, or a template several servers'ModConfigfolders are built from, gives all of them one id, and two servers with the sameServiceNameand the same id are one server to a backend. Clear the key in the copies, or give each server its own value. - Pulse OTLP cannot save the generated id when
ModConfigis read-only. It then logs a warning naming the id and saying the next start will export a different one. PutServiceInstanceIdin the file yourself, or setOTEL_RESOURCE_ATTRIBUTES=service.instance.id=<id>in the server's environment. AModConfigthat does not survive a restart (a container without a volume for it) loses the id the same way, with no warning. - The first start rewrites
pulse-otlp.jsonto add the key. Like every config upgrade, the rewrite drops comments in the file, and names any key Pulse does not know in a warning before dropping it.
The environment still has the last word. A service.instance.id in OTEL_RESOURCE_ATTRIBUTES wins over the key, and it now survives on its own, where 0.2.0 replaced it on every start unless OTEL_SERVICE_NAME was set too. With OTEL_SERVICE_NAME set, nothing changes: the environment decides the whole identity, as before.
Attribution: the game's own work no longer lands in unattributed
This one only matters with attribution on. Part of the game's own entity behaviours (despawn timers, name tags, creature physics and more) was filed under unattributed, because the name a behaviour marks the tick with is often not the code its class was registered under. Pulse now reads the loaded entities to learn which mod ships each behaviour, and that time goes to game, survival or engine. On a test server with 1,691 entities, unattributed went from about 2.5% of the sampled tick to under 0.1%.
The read happens outside the profiled ticks. The first burst after attribution starts reads every loaded entity: about 3 to 4 ms with 1,700 entities, 7 ms with 4,000, 23 to 40 ms with 8,000. After that a burst costs about 0.1 ms, plus the entities that loaded in between.
Documentation
The getting-started guide now covers macOS, gives the right data path for the server.sh that ships with the server (/var/vintagestory/data/Mods), and explains how to load the bundled alert rules into Prometheus.
Upgrading
Replace pulse_0.2.0.zip in Mods/ with pulse_0.2.1-indev.1.zip, do the same for pulseotlp if you push OTLP, and restart. This is a prerelease: try it on a server you can restart. The stable 0.2.1 follows once it has run on a real one.
Full changelog: v0.2.0...v0.2.1-indev.1
Pulse v0.2.0
The first stable release since 0.1.0. Wake up, admin: your server has a pulse, and now it can tell you which mod is spending it.
This tag folds in everything published across the five 0.2.0 prereleases (indev.1 through indev.5), plus a round of fixes that landed after indev.5. If you already run indev.5, read "Fixed since indev.5" below. If you are upgrading from 0.1.0, read the whole thing: there is one security note and three upgrade hazards that need your attention, a Bind: 0.0.0.0 setting that finally takes effect, a breaking metric rename and, if you push OTLP, a change to the job and instance labels your backend assigns each series.
Security note: Docker commands in earlier READMEs
If you ever ran the docker run commands from an earlier version of contrib/grafana/README.md or contrib/alerts/README.md, read this first. Those commands started Grafana and Prometheus with host networking and no bind to loopback, and Grafana with anonymous admin access on top. On Linux, that Grafana gives administrator access to anyone who can reach port 3000 on the machine, and Prometheus serves your server's metrics on port 9090 the same way. Stop both containers (docker stop pulse-graf pulse-prom), then start them again from an up-to-date copy of the kit: docker compose up -d in contrib/grafana, or, if you load the alert rules, the current docker run commands in contrib/grafana/README.md and contrib/alerts/README.md. All of them bind to 127.0.0.1 only. To view a remote server's dashboard, use an SSH tunnel; docs/getting-started.md shows how.
Upgrading from 0.1.0
Replace pulse_0.1.0.zip in Mods/ with pulse_0.2.0.zip, do the same for pulseotlp if you push OTLP, and restart once. Nothing needs hand-editing: each mod checks its config file at startup and writes back any key it is missing, with that key's default, so pulse.json picks up the new Attribution block and pulse-otlp.json picks up ServiceName on their own. Values you already set are left exactly as they are. A config key neither mod recognises is named in a warning instead of disappearing quietly, since the rewrite drops it. That rewrite only happens when a key is missing, and it writes a fresh copy of the file from the loaded values, so a comment you hand-added to either file will not survive one.
If your pulse.json still sets Bind to 0.0.0.0 from an earlier attempt, check it before you restart: that value failed to bind under 0.1.0 (Pulse logged an error at every start and served nothing), so an admin who set it and moved on may never have noticed. The plain-socket endpoint (see What's new) binds it, so this restart is the one where 0.0.0.0 starts working, serving /metrics on every interface with no authentication, exactly what 0.1.0's own README warned against.
Every 0.2.0 start also logs the engine's own warning, "Over 400ms tick. Skipping N physics ticks.", once at startup, since Pulse primes the engine's frame profiler for the first tick after every start (attribution on or off) and the engine only prints that line while its profiler is on; it adds 1 to pulse_log_entries_total{level="warning"} and is otherwise harmless, since it is not one of the kinds pulse_engine_warnings_total counts, so no bundled alert fires over it. With attribution on, the same line also appears during a profiled burst whenever physics falls behind.
If you push OTLP, this upgrade also changes the service.name your backend sees, and with it the job label Prometheus's OTLP receiver, Mimir and Grafana Cloud derive from it, plus an instance label 0.1.0's series never carried. 0.1.0 never set service.name itself, so the OpenTelemetry SDK exported its own fallback, unknown_service: followed by the process name (dotnet when launched as dotnet VintagestoryServer.dll, VintagestoryServer through the native apphost), unless you had already set OTEL_SERVICE_NAME or OTEL_RESOURCE_ATTRIBUTES=service.name=...: the SDK's own default resource already reads both, with no ServiceName key involved and no instance label.
Unless OTEL_SERVICE_NAME is set, 0.2.0 calls AddService, and two things follow. Left at its default, ServiceName changes every existing series' job to vintagestory, so a dashboard variable or alert keyed on the old value breaks quietly rather than erroring. Whatever ServiceName holds, every series also gains an instance label, from a service.instance.id that AddService regenerates at random on every restart.
Setting ServiceName to your old value brings the old job back but not the old identity: you still get the new, restart-churning instance label. To keep 0.1.0's job and instance labels exactly, set OTEL_SERVICE_NAME to your old value instead: the mod then skips AddService entirely, so neither job nor instance changes. If you already had OTEL_SERVICE_NAME set under 0.1.0, neither job nor instance changes either, though target_info's telemetry_sdk_version moves from 1.18.0 to 1.19.1 and the resource gains a schema URL.
A service.name or service.instance.id set through OTEL_RESOURCE_ATTRIBUTES is silently overridden by AddService unless OTEL_SERVICE_NAME is also set. Set OTEL_SERVICE_NAME to move your service.name there; the service.instance.id from OTEL_RESOURCE_ATTRIBUTES then survives, since that also skips AddService.
For a stable instance label across restarts, set both: OTEL_SERVICE_NAME=<name> together with OTEL_RESOURCE_ATTRIBUTES=service.instance.id=<id>. OTEL_RESOURCE_ATTRIBUTES alone is not enough, since AddService's own randomly generated id silently overrides it on every start.
If you ever roll back to 0.1.0, it leaves those config files exactly as they are: it does not know Attribution or ServiceName and does not strip them, it just ignores them, so nothing on disk is lost, but any attribution tuning goes inert until you upgrade again, and so does ServiceName: 0.1.0 ignores that key and goes back to the SDK's own unknown_service: fallback job, unless OTEL_SERVICE_NAME is set. 0.1.0 honours that variable through the same SDK default resource 0.2.0 does, so a name set that way carries through the rollback unchanged; a name set only through the ServiceName config key does not.
What's new
- Per-mod tick attribution. Turn it on and see which mod is actually spending the tick, on a live graph rather than a one-off report:
pulse_mod_tick_share{modid},pulse_mod_tick_seconds_total{modid}, and two counters behind them. Off by default; see the cost note below before you turn it on. /pulse, a new server command (behindcontrolserver)./pulse attribution on|off|statusswitches the duty cycle on a running server with no restart, without writingpulse.json, so a switch does not survive one, and/pulse reloadre-readspulse.jsonand applies theAttributionblock live, naming any other key that still needs a restart to take effect.- The metrics endpoint is now a plain socket, not
HttpListener. Binding loopback on Windows no longer needs administrator rights or anetshreservation, a scrape ofhttp://localhost:9464/metricson Linux no longer gets a 404 because its Host header does not matchBind, several scrapes are served at once instead of queued, andBind: 0.0.0.0now binds instead of failing with a logged error. - Config files upgrade and tolerate damage. Covered above for the happy path; a
pulse.jsonorpulse-otlp.jsonthat exists but will not parse (a doubled comma, a missing quote) no longer stops the mod from starting. Each mod logs the file's path and the parser's own message and leaves the file untouched; Pulse runs that session on its built-in defaults, and Pulse OTLP leaves export off for that session rather than guess an endpoint. - OTLP export failures no longer pass silently. A rejected push, a refused or unreachable collector, a timeout: each kind now logs one line, repeated at most every 10 minutes, naming what failed and the backend's own (redacted) response. A matching line reports the first successful export after a failure, or the first export after each server start when nothing has failed yet. Six families that could stay invisible on a quiet server until their first real event (engine warnings, player deaths, suspends, suspend seconds, worldgen columns, log entries) now reach OTLP seeded at zero from the first export instead of waiting. A new
ServiceNameconfig key (defaultvintagestory) sets theservice.nameresource attribute; set it per server so a backend collecting from more than one can tell them apart, since the default is the same on every install.OTEL_SERVICE_NAMEstill overrides it if you set that instead. contrib/alerts, a new Prometheus alerting rules pack, eleven rules covering tick rate, tick saturation, sustained tick overruns, engine warnings, log errors, endpoint availability and a stuck worldgen queue, one of themPulseModHoggingTick, which only fires when a single mod holds more than half the profiled tick for 10 minutes while the server is also over 80% of its tick budget. The bundled Grafana dashboard gains a matching attribution row.docs/getting-started.md, a walkthrough for a server owner who has never used Prometheus or Grafana, routed by how the server is hosted, ending at the shared dashboard either way.
Breaking change: nine runtime series renamed
Prometheus, Mimir and Grafana Cloud derive a name from an OTLP instrument by looking at its unit. Nine dotnet_* families on /metrics used to skip that step, only mapping dots to underscores and appending _total, so they were out of step with every other OTLP-derived name on the same dashboard. From 0.2.0 they match:
| Old name | New name |
|---|---|
dotnet_process_memory_working_set |
dotnet_process_memory_working_set_bytes |
dotnet_gc_heap_total_allocated_total |
dotnet_gc_heap_allocated_bytes_total |
| ... |
Pulse v0.2.0-indev.5
Everything merged since v0.2.0-indev.4. Three points need your attention before or after upgrading; the rest works as it did.
If you followed the Docker commands from an earlier README
The docker run commands in earlier versions of contrib/grafana/README.md and contrib/alerts/README.md started Grafana and Prometheus with host networking, without binding them to loopback, and Grafana with anonymous admin access. That Grafana answers on every interface of the machine, to anyone who can reach port 3000, as an administrator, and Prometheus serves your server's metrics on port 9090 the same way. Stop both containers (docker stop pulse-graf pulse-prom) and start them again with the new commands or the new docker-compose.yml, which bind both to 127.0.0.1 only. To view a remote server's dashboard, use an SSH tunnel; the guide shows how.
Nine runtime series have new names
They now match the standard names Prometheus and Grafana derive from the same .NET instruments over OTLP, so one name serves both paths. If a panel or alert of yours queries one of them, update it:
| Before | From this version |
|---|---|
dotnet_process_memory_working_set |
dotnet_process_memory_working_set_bytes |
dotnet_gc_heap_total_allocated_total |
dotnet_gc_heap_allocated_bytes_total |
dotnet_gc_last_collection_memory_committed_size |
dotnet_gc_last_collection_memory_committed_size_bytes |
dotnet_gc_last_collection_heap_size |
dotnet_gc_last_collection_heap_size_bytes |
dotnet_gc_last_collection_heap_fragmentation_size |
dotnet_gc_last_collection_heap_fragmentation_size_bytes |
dotnet_gc_pause_time_total |
dotnet_gc_pause_time_seconds_total |
dotnet_jit_compiled_il_size_total |
dotnet_jit_compiled_il_size_bytes_total |
dotnet_jit_compilation_time_total |
dotnet_jit_compilation_time_seconds_total |
dotnet_process_cpu_time_total |
dotnet_process_cpu_time_seconds_total |
No pulse_* series changed. The bundled dashboard reads both names for the whole 0.2 line, so it keeps working while you upgrade servers one at a time.
Attribution costs more than earlier notes said
The notes for indev.2 to indev.4 quoted about 0.3% of the tick budget, amortised. That was an estimate from counting profiler marks, and it was too low. Measured on a server with 4000 entities, a profiled tick costs about 26% of the budget. At the new default burst of 10 ticks every 10 seconds, that comes to about 0.9%; the old default of 30 ticks comes to about 2.5%. A pulse.json written by an earlier version keeps its BurstTicks of 30: set it to 10 and run /pulse reload to get the new default. Attribution is still off unless you turn it on.
Metrics endpoint
It is now served from a plain socket instead of .NET's HttpListener:
http://localhost:9464/metricsworks, whatever Host header the client sends. Linux and macOS used to answer 404 to anything but the exact bind address.- On Windows, no administrator rights or
netshreservation are needed. Bindset to0.0.0.0works on Linux.- Several scrapes are served at once, and each connection has a hard 5 second deadline, so one stalled client no longer holds up the others.
- If you widen
Bindbeyond loopback, add a firewall rule limiting the port to your scraper: the endpoint has no authentication, and anyone who can reach it can occupy its connection slots.
Pulse OTLP
- Export failures are logged now: one Warning per kind of failure, repeated at most every 10 minutes, including the backend's answer when there is one (a wrong Grafana Cloud token shows the 401 and its message). Configured header values and anything shaped like a bearer or basic credential are masked in that line, but treat the log as sensitive when you share it. The first successful export logs one line too, and so does the first one after a failure, so you can see it working and see it recover.
- Five counter families (engine warnings, player deaths, suspends, suspend seconds, worldgen columns) no longer stay missing from an OTLP-fed dashboard on an idle server.
- An
Endpointwith a query string keeps it; the metrics path used to be appended after the query. The startup line no longer prints credentials or the query string. - OpenTelemetry .NET 1.19.1.
Config and attribution
- A
pulse.jsonorpulse-otlp.jsonthat does not parse (a doubled comma, a missing quote) no longer stops the mod. It logs the file's path and the parser's message, leaves the file as it is, and runs on built-in defaults for that session. - The per-mod share no longer keeps showing the last burst's values once attribution stops.
- The dashboard has an attribution row, and
contrib/alertshas a rule for a mod taking more than half of the busy tick time for 10 minutes while the server runs above 80% of its tick budget.
Getting started
docs/getting-started.md walks a server owner from installing the mod to a first dashboard, either locally with Prometheus and Grafana or on Grafana Cloud over OTLP.
Same game support: Vintage Story 1.22.x, dedicated server, .NET 10. Test bar for this tag: 407 tests including 39 embedded-server scenarios, 79 mutations killed, quality gate green, and an end-to-end run of both release zips on a Vintage Story 1.22.7 dedicated server with the bundled Prometheus, Grafana, alert rules and an OTLP receiver.
Pulse v0.2.0-indev.4
Per-mod tick attribution no longer needs a restart. That mattered: the moment you want a profiler is while the server is struggling, and a restart throws away exactly what you wanted to look at.
A new server command, /pulse, drives it live (privilege controlserver, so a panel console can use it too):
/pulse attribution onstarts the duty cycle now,/pulse attribution offstops it and leaves the engine profiler switched off. Both act on the running server only and never writepulse.json, so a ten-minute look does not become permanent by accident./pulse attribution statustells you what is running: on or off, burst and interval in use, ticks profiled so far./pulse reloadre-readspulse.jsonand applies the wholeAttributionblock live. Keys that cannot change at runtime (Port,Bind,RuntimeMetrics,ChunksRefreshSeconds) are named in the reply when their value differs, so you know a restart is still owed.
To make the live switch safe, the attribution families are now always registered and the engine's frame profiler is always primed at startup, whether or not Attribution.Enabled is set. An idle server exposes no extra series: the families only emit once a burst has completed.
Everything else is unchanged from v0.2.0-indev.3: same metric families, same names, same labels. Both zips are published as usual; the base mod is still one dll with no dependencies.
Same game support: Vintage Story 1.22.x, dedicated server, .NET 10. Test bar for this tag: 217 tests including 30 embedded-server scenarios (two of them run the /pulse commands on a live server and check the profiler is left switched off), 44 mutations killed, quality gate green.
Pulse v0.2.0-indev.3
A small prerelease for a real production annoyance: a pulse.json written by an older version never showed the keys a newer version added, so an admin upgrading from 0.1.0 had no way to discover the Attribution block short of reading the README, and the same went for ServiceName in pulse-otlp.json.
From this build, both mods complete their config file at startup:
- Keys missing from the existing file are added with their defaults, and the server log lists which ones. Your own values are kept exactly as they were.
- Keys the mod does not know (a typo, a retired option) are dropped, with a warning naming them, so a mistake never disappears silently.
- A file that is already complete is not touched at all: no write, no modification time change, which matters on hosts that mount configs read-only or keep them under version control.
Nothing else changes: same metric families, same names, same labels as v0.2.0-indev.2. Both zips are published as usual; the base mod is still one dll with no dependencies.
Same game support: Vintage Story 1.22.x, dedicated server, .NET 10. Test bar for this tag: 195 tests including 28 embedded-server scenarios (two of them seed a partial config with a typo in it and read the upgraded file back off disk), 41 mutations killed, quality gate green.
Pulse v0.2.0-indev.2
This prerelease answers the question every busy server eventually asks: which mod is eating the tick.
Turn on the new Attribution block in ModConfig/pulse.json and Pulse drives the engine's own frame profiler in short bursts (30 ticks every 10 seconds by default), reads the marks it stamps on every tick listener, delayed callback and main-thread entity behavior, and maps each one back to the mod that registered it. Three new families come out of it, plus a meta counter:
pulse_mod_tick_share{modid}: each mod's fraction of the profiled busy time over the last burst. The shares sum to 1, so they read straight off a dashboard.pulse_mod_tick_seconds_total{modid}: profiled seconds attributed per mod, cumulative. Sampled during bursts, so treat it as a ratio source, not a wall-clock total.pulse_attribution_ticks_total: how many ticks were actually profiled, for normalising.pulse_attribution_dropped_samples_total: marks discarded because their elapsed time overflowed an engine-side 32 bit counter.
Off by default, and priced before shipping: a burst costs a few percent of the tick budget while it runs (scaling with loaded entities), which the default duty cycle amortises to roughly 0.3%. The engine's profiler has a genuine crash trap when enabled mid-tick on a cold server; Pulse primes it before the tick loop exists and refuses to start a burst until the profiler has proven itself warm, so the trap is structurally unreachable. Blind spots are documented in the README: broadcast events (PlayerJoin and friends) carry no marks at all, and threaded physics is attributed for its main-thread slice only.
Everything lands in the base pulse zip, still a single dll with no dependencies. The OTLP companion is unchanged and pushes the new families like any others.
Same game support: Vintage Story 1.22.x, dedicated server, .NET 10. Test bar for this tag: 183 tests including 26 embedded-server scenarios (one boots a server, runs a burst and checks the shares sum to 1), 38 mutations killed, quality gate green.
Pulse v0.2.0-indev.1
First prerelease of the 0.2 line, carrying everything that landed since 0.1.0 went stable. Nothing here changes the metrics themselves; this wave is about running Pulse in production and trusting what it sends.
What's new:
- An operations kit:
contrib/alerts/pulse-alerts.ymlships ready-made Prometheus alerting rules for tick rate, tick saturation, sustained overruns, engine warnings, log errors, endpoint availability and a stuck worldgen queue, calibrated against the engine's own thresholds. Load it next to the provisioned Grafana dashboard and the monitoring side is done. ServiceNameinpulse-otlp.json(defaultvintagestory): sets theservice.nameresource attribute, so a backend receiving several servers can tell them apart.OTEL_SERVICE_NAMEstill wins when set, as the ecosystem expects.- A weekly CI tripwire now builds and tests the mod against the newest stable Vintage Story version, so an engine update that breaks the deep-engine probe is caught on the game's schedule rather than by a server admin.
- The grpc setting in
pulse-otlp.jsonis now proven on the wire: a scenario boots a real server against a fake gRPC collector and reads the export off the socket. Both protocols the config accepts are exercised end to end.
Two zips as usual: pulse_0.2.0-indev.1.zip alone gives you the Prometheus endpoint; add pulseotlp_0.2.0-indev.1.zip next to it in Mods/ to push over OTLP. Same game support: Vintage Story 1.22.x, dedicated server, .NET 10. Test bar for this tag: 152 tests including 21 embedded-server scenarios, 32 mutations killed, quality gate green.
Pulse v0.1.0
The first stable release of Pulse: your Vintage Story server's vital signs on a Prometheus scrape endpoint.
Drop pulse_0.1.0.zip into the server's Mods/ folder, start the server, and http://127.0.0.1:9464/metrics answers with the full picture: tick rate and tick time (histogram, p95 and p99 included) against the configured budget, the engine's own tick busy time (the /stats number, the one that shows load climbing long before TPS drops), players with ping and deaths, entities with a top-ten breakdown by type, chunks, the worldgen queue, TCP and UDP traffic, autosave pause counts and durations, log health counters with four engine warnings worth alerting on, and the .NET runtime's own GC, memory, CPU and thread pool families.
Three ways to consume it: point a Prometheus at the endpoint (the contrib/grafana folder holds a fully provisioned dashboard behind two docker commands), let a game panel read /metrics directly (the first panel integration was built that way in an afternoon), or add pulseotlp_0.1.0.zip beside the base mod to push everything over OTLP to an OpenTelemetry collector, Grafana Cloud, Datadog and the rest.
The base mod is one dll with nothing bundled. The endpoint binds loopback by default, on purpose. A bind failure, an unreachable collector or a future engine change in the deep-metrics probe all degrade cleanly and can never take the game server down.
Field record for this release line: five prereleases in the open, a hosting provider running it in production testing, a third-party panel consuming the endpoint, and the OTLP path verified against the official collector. Ships from CI with 126 unit tests, 20 embedded-server scenarios and a 31-mutation check, quality gate green at the 80% coverage bar.
Vintage Story 1.22.x, dedicated server only, .NET 10. Also runs on the Stratum and Lithos server forks.
Pulse v0.1.0-indev.5
The stable candidate. Nothing new to install for the sake of features; this build settles the last two metric names before the 0.1.0 release freezes them as a contract, and it ships behind the day's real-world validation.
Renamed, if you already built panels on the indev.4 names:
pulse_server_tick_busy_seconds_avgis nowpulse_server_tick_busy_seconds.pulse_network_packets_in_window{channel}andpulse_network_bytes_in_window{channel}are nowpulse_network_packets_per_second{channel}andpulse_network_bytes_per_second{channel}, and the values are rates: the engine's completed window divided by its nominal two seconds. A dashboard wants a rate, and the name no longer leaks an engine internal.
Everything else is identical to v0.1.0-indev.4. Since that tag, the OTLP path was exercised against the official OpenTelemetry collector receiving live batches from a standalone server, on top of the existing fake-collector scenario.
Unless testing turns something up, 0.1.0 stable will be this build with the version stamp and no other change. Test bar: 126 unit tests, 20 embedded-server scenarios, 31 mutations, all green in CI from the tag.