Repository navigation
Monitoring
Every provider can serve Prometheus metrics at /metrics. The monitoring/ bundle runs Prometheus and Grafana with a ready-made dashboard, so you can watch a whole fleet from one page.
It takes two steps: turn metrics on at each provider, then start the bundle on any machine that can reach them.
Prometheus is the fleet view: it answers "how is this node doing right now", across every box, on a live dashboard. It is also the heavier option, and it says nothing about how a box behaved before an upgrade unless you were already recording.
Every provider therefore also keeps a small local record of its own behaviour in ~/.urnetwork/baseline.jsonl, on by default, with no setup. It holds counts and totals only, never proxy addresses, usernames or passwords. The same file also records the GC state (gc: GOGC in force, whether the governor is tightening, and the GC CPU share) and the kernel's socket-table pressure (net: conntrack fill, TIME_WAIT and orphaned TCP sockets). These are observation only and feed no decision; a field the box cannot read is left out rather than written as zero. urnet-tools baseline show prints the newest rows and urnet-tools baseline compare reports a rate, proxy count, memory and restart difference across an upgrade, warning you when the capacity changed so a trimmed box is not misread as a regression.
This is the no-setup alternative for the one question Prometheus does not answer: did that upgrade make this box worse? It is not a replacement for the fleet view, and it is not a time-series database. For a fleet, use the bundle below.
urnet-tools metrics onThe command prints where the endpoint listens and what to give Prometheus:
Metrics: on
listening http://127.0.0.1:9100/metrics
listening http://100.64.0.10:9100/metrics
Prometheus target: 100.64.0.10:9100
The setting survives restarts and updates. Run urnet-tools metrics status at any time to see it again.
| Situation | Listens on |
|---|---|
| Bare metal or VM (default) |
127.0.0.1 plus this machine's Tailscale address, if it has one |
| Docker container (default) | Every interface inside the container |
urnet-tools metrics listen <ip:port> |
Exactly that address |
URNETWORK_METRICS=<ip:port> set in the environment |
Exactly that address (overrides everything above) |
The default port is 9100. If another program already uses it (Prometheus's node_exporter does), the provider takes the next free port up to 9103. Use whatever metrics status prints.
Tip
Tailscale is the easiest way to connect Prometheus to providers on different networks. If Tailscale starts after the provider, the provider starts serving on the Tailscale address within 30 seconds. No restart is needed.
Without Tailscale, pick an address Prometheus can reach, such as a private LAN or VPN address:
urnet-tools metrics listen 10.0.0.5:9100
urnet-tools metrics listen auto # back to the defaultWarning
The metrics include every proxy's address and traffic. Never expose the endpoint on a public address. metrics listen 0.0.0.0:9100 on an internet-facing server is readable by anyone who can reach that port, unless a firewall blocks it.
Publish the port on an address Prometheus can reach, then turn metrics on inside the container:
docker run -d --name urnetwork -p 100.64.0.10:9100:9100 ... ghcr.io/full-bars/meso-miner:latest
docker exec urnetwork urnet-tools metrics onHere 100.64.0.10 is the host's Tailscale address, and the Prometheus target is 100.64.0.10:9100. For several containers on one host, publish each on its own host port (-p 100.64.0.10:9101:9100) and use that port as the target.
Warning
With --network host, the container shares the host's interfaces, so the default binds every host interface, including public ones. Use urnet-tools metrics listen <tailscale-ip>:9100 there.
You need Docker with Compose v2 on the monitoring machine. Download urnetwork-monitoring-<version>.tar.gz from the latest release, then:
tar xzf urnetwork-monitoring-*.tar.gz
cd monitoring
./setup.sh 100.64.0.10:9100=node-1 100.64.0.11:9101=node-2
docker compose up -dEach address=name argument adds one provider. Use the address from metrics status; the name is what the dashboard shows. On first run, setup.sh also creates .env with a random Grafana admin password.
Open http://<monitoring machine>:3000 and sign in as admin with the password from .env. The URnetwork Providers dashboard is the home page.
To add providers later, run ./setup.sh again with the new ones. Prometheus picks them up within a minute, with no restart.
Important
Grafana listens on port 3000 on every interface, protected by the admin password. To keep it private, set GRAFANA_BIND in .env to 127.0.0.1 or the machine's Tailscale address, then run docker compose up -d again.
Add any of these to .env, then run docker compose up -d:
| Variable | Default | Meaning |
|---|---|---|
GRAFANA_BIND |
0.0.0.0 |
Address Grafana listens on |
GRAFANA_PORT |
3000 |
Grafana port |
PROMETHEUS_PORT |
9090 |
Prometheus port, always on 127.0.0.1
|
PROMETHEUS_RETENTION |
30d |
How long Prometheus keeps data |
- Fleet: nodes up and down, billable throughput, active clients, and proxies up or failing. A table lists each node's version, uptime, clients, proxies, memory, and resource pressure.
- Traffic: billable and total bytes per second, clients per node, and the top 15 proxies by billable traffic.
- Proxies: pool status, grades, recoveries and losses, and contract outcomes.
- Node health: errors by category, resource pressure, memory, goroutines, sessions, and DNS-over-HTTPS failures. Two panels plot memory and file descriptors against their limits (the limit is a dashed line).
- Transport (H1 and H3): payload bytes per second by transport mode and direction, the share of outbound frames carried by H3, H3 connections up, H3 connect attempts, failures and drops per hour, how many connections offered and had QUIC DATAGRAM accepted, datagram messages received and dropped per second, and the WhoDis DNS and DNS-pump transports (connections up, attempts and failures). Auto mode does not start the DNS modes, so that panel stays at zero unless a transport was built with one of those target modes.
- Lifecycle: restarts by reason, uptime per node, HotSwap outcomes per hour, and version skew (each version in the fleet and how many nodes run it).
The transport panels need a provider that exports the urnet_transport_* and urnet_h3_* series. On older providers they stay empty, and the H3 and DATAGRAM series only move once h3 and h3_datagram are switched on.
Use the Node selector at the top to narrow the view to specific providers.
The memory, descriptor and restart-reason panels need a provider that exports urnet_mem_limit_bytes, urnet_rss_bytes, urnet_open_fds, urnet_fd_limit and urnet_restart_reason. On older providers those panels stay empty and the rest of the dashboard works as before. Descriptor metrics are Linux only.
The bundle ships alert rules in monitoring/prometheus/rules/urnetwork.yml. Prometheus loads them at startup, and docker compose up -d mounts the rules directory for you. Firing alerts show at http://127.0.0.1:9090/alerts.
Note
The bundle has no Alertmanager. Alerts are visible in Prometheus and as the ALERTS series, but nothing is emailed or posted until you add an Alertmanager and point Prometheus at it.
| Alert | Fires when | Does not detect |
|---|---|---|
UrnetworkNodeDown |
A node's /metrics endpoint has not answered for 5 minutes (up == 0). |
A stopped provider looks the same as metrics switched off, a blocked port or a Tailscale outage. A node removed from targets/providers.yml never fires. It says nothing about whether a reachable node is earning. |
UrnetworkRestartLoop |
More than 3 restarts of one node in an hour. | A restart that hides between two scrapes, and the cause of a restart. Check urnet_restart_reason and the node's logs. |
UrnetworkRestartLoop counts drops of urnet_uptime_seconds with resets(). It does not use urnet_startup_restarted, because that gauge keeps one fixed value for the life of a process, so it cannot be counted with increase().
These are in the same file, commented out, because they need thresholds that suit your fleet.
| Alert | Fires when | Does not detect |
|---|---|---|
UrnetworkOldVersion |
A node runs a different version from most of the fleet for 7 days. | Which version is newer. During a slow rollout it can flag the early upgraders. It does not fire on a fleet that is uniformly out of date. |
UrnetworkMemoryNearLimit |
urnet_mem_sys_bytes is above 90% of urnet_mem_limit_bytes for 15 minutes. |
Resident memory, other processes on the machine, and a container limit lower than the Go limit. Nodes with no limit set are skipped. |
UrnetworkFileDescriptorsNearLimit |
urnet_open_fds is above 85% of urnet_fd_limit for 10 minutes. |
Which kind of descriptor leaks, and a spike between two scrapes. |
UrnetworkNoBillableTraffic |
A node is up, has been up for over 30 minutes, and billed no traffic for 30 minutes. | A broken node versus a quiet one with no demand. It can be noisy on idle nodes. |
To enable one, open monitoring/prometheus/rules/urnetwork.yml, read the note above the rule, adjust the threshold, and remove the leading # from every line of that block (the block ends at the next blank line). Then reload Prometheus:
curl -X POST http://127.0.0.1:9090/-/reloadNodes whose provider does not export a metric never match its rule, so a fleet with mixed versions is safe.
If you have promtool (it ships with Prometheus), run these from the monitoring/ directory after editing the rules:
promtool check rules prometheus/rules/urnetwork.yml
promtool test rules prometheus/rules/urnetwork.test.ymlThe test file covers the two enabled rules. prometheus.yml lists urnetwork.yml by name, so the test file is never loaded as rules.
A node shows as down. On that node, run urnet-tools metrics status and check that the address matches the target in prometheus/targets/providers.yml. From the monitoring machine, curl http://<target>/metrics | head should print metrics. connection refused means nothing listens on that address. A timeout usually means a firewall or Tailscale ACL is blocking the port.
The target in Prometheus shows an error. Open http://127.0.0.1:9090/targets on the monitoring machine. The error column says why the last scrape failed.
Panels are empty right after starting. Rates need two scrapes, so allow about a minute.
| Metric | Type | Labels | Meaning |
|---|---|---|---|
urnet_info |
gauge | version |
Always 1; carries the provider version |
urnet_uptime_seconds |
gauge | Provider uptime | |
urnet_billable_bytes_total |
counter | direction |
Billable bytes, all proxies |
urnet_bytes_total |
counter | direction |
All bytes, all proxies |
urnet_clients_active |
gauge | Active clients, all proxies | |
urnet_connections_active |
gauge | Active connections | |
urnet_transport_frames_total |
counter |
mode, dir
|
Payload frames carried by each platform transport mode (h1, h3, h3dns, h3dnspump) per direction. Keepalives and speed or latency echoes are excluded |
urnet_transport_payload_bytes_total |
counter |
mode, dir
|
Payload bytes carried by each platform transport mode, before framing |
urnet_transport_direct_h1_frames_total |
counter | dir |
Payload frames carried by H1 for the direct identity only. This is the fair comparison for H3, which runs for that identity alone; urnet_transport_frames_total{mode="h1"} is every proxy identity together |
urnet_transport_direct_h1_payload_bytes_total |
counter | dir |
Payload bytes carried by H1 for the direct identity only |
urnet_h3_up |
gauge | H3 connections up now | |
urnet_h3_connect_attempts_total |
counter | H3 connect attempts | |
urnet_h3_connects_total |
counter | H3 connections that authenticated | |
urnet_h3_connect_failures_total |
counter | H3 connect attempts that failed | |
urnet_h3_drops_total |
counter | H3 connections that ended while still wanted. Switching h3 or h3_datagram is not one |
|
urnet_h3_datagram_offered_total |
counter | H3 connections that offered QUIC DATAGRAM | |
urnet_h3_datagram_accepted_total |
counter | H3 connections where the server accepted DATAGRAM. Offered without accepted means an old server | |
urnet_h3_datagram_rx_messages_total |
counter | Messages received over DATAGRAM | |
urnet_h3_datagram_rx_bytes_total |
counter | Bytes of messages received over DATAGRAM | |
urnet_h3_datagram_rx_dropped_total |
counter | Received DATAGRAM messages dropped because the receive route was full | |
urnet_h3_datagram_rx_rejected_total |
counter | reason |
DATAGRAMs refused by the datagram layer: malformed, duplicate, checksum
|
urnet_h3_datagram_tx_messages_total |
counter | Messages sent over DATAGRAM | |
urnet_h3_datagram_tx_bytes_total |
counter | Bytes of messages sent over DATAGRAM | |
urnet_h3_datagram_tx_stream_messages_total |
counter | Messages sent on the reliable stream of a connection that negotiated DATAGRAM, so the lane split is tx_messages against this | |
urnet_h3_datagram_tx_errors_total |
counter | DATAGRAM send errors | |
urnet_h3_datagram_blackholes_total |
counter | Connections whose datagram send was switched off because datagrams went out and none came back | |
urnet_transport_pt_up |
gauge | mode |
Packet-translation (h3dns, h3dnspump) transport connections up now |
urnet_transport_pt_connect_attempts_total |
counter | mode |
Packet-translation transport connect attempts |
urnet_transport_pt_connects_total |
counter | mode |
Packet-translation transport connections that authenticated |
urnet_transport_pt_connect_failures_total |
counter | mode |
Packet-translation transport connect attempts that failed |
urnet_transport_pt_drops_total |
counter | mode |
Packet-translation transport connections that ended while still wanted |
urnet_proxy_pool_size |
gauge | status |
Proxies by status: up, connecting, degraded, dead |
urnet_proxy_bytes_total |
counter |
proxy, direction
|
Bytes per proxy |
urnet_proxy_billable_bytes_total |
counter |
proxy, direction
|
Billable bytes per proxy |
urnet_proxy_clients |
gauge | proxy |
Active clients per proxy |
urnet_proxy_session_age_seconds |
gauge | proxy |
Length of the current client presence window per proxy |
urnet_proxies_known |
gauge | Proxies in the provider's proxy state | |
urnet_proxy_grades |
gauge | tier |
Proxies by grade |
urnet_proxy_health |
gauge | status |
Proxies by recorded health |
urnet_proxy_graded_recent |
gauge | Proxies graded in the last hour | |
urnet_proxy_graded_stale |
gauge | Proxies with an older grade | |
urnet_proxy_auth_failures |
gauge | Auth failures summed over current proxies | |
urnet_url_proxy_grades |
gauge | tier |
URL-sourced proxies by grade |
urnet_url_proxy_ungraded |
gauge | URL-sourced proxies not yet graded | |
urnet_contracts_total |
counter | result |
Contract outcomes since start |
urnet_errors_total |
counter | category |
Errors since start |
urnet_doh_failures_total |
counter | DNS-over-HTTPS failures | |
urnet_sessions_pqe |
gauge | Active post-quantum sessions | |
urnet_sessions_classical |
gauge | Active classical sessions | |
urnet_sessions_opened_total |
counter | encryption |
Sessions opened, lifetime |
urnet_sessions_opened_recent |
gauge |
window, encryption
|
Sessions opened in the last hour, day, or week |
urnet_lifetime_sessions_total |
counter | encryption |
Sessions, persisted across restarts |
urnet_lifetime_contracts_total |
counter | result |
Contract outcomes, persisted across restarts |
urnet_lifetime_proxies_total |
counter | event |
Proxy recoveries and losses, persisted across restarts |
urnet_lifetime_billable_bytes_total |
counter | Billable bytes, persisted across restarts | |
urnet_lifetime_errors_total |
counter | category |
Errors, persisted across restarts |
urnet_hotswap_outcomes_total |
counter | reason |
HotSwap attempt outcomes |
urnet_control_commands_total |
counter | cmd |
Control socket commands |
urnet_startup_clean_shutdown |
gauge | 1 if the previous run shut down cleanly | |
urnet_startup_restarted |
gauge | 1 if the previous run did not shut down cleanly | |
urnet_startup_upgraded |
gauge | 1 if the version changed since the previous run | |
urnet_startup_previous_version |
gauge | version |
Always 1; the version before the last upgrade |
urnet_pressure_score |
gauge | Resource pressure, 0 (fine) to 1 (emergency) | |
urnet_mem_heap_bytes |
gauge | Go heap in use | |
urnet_mem_sys_bytes |
gauge | Memory obtained from the OS | |
urnet_mem_limit_bytes |
gauge | The Go memory limit in effect | |
urnet_rss_bytes |
gauge | Resident set size (Linux) | |
urnet_open_fds |
gauge | Open file descriptors (Linux) | |
urnet_fd_limit |
gauge | File descriptor limit (Linux) | |
urnet_restart_reason |
gauge | reason |
Value 1 for the reason the current process started; absent for other reasons |
urnet_gc_cycles_total |
counter | Garbage collection cycles | |
urnet_gc_gogc |
gauge | GOGC value currently in force | |
urnet_gc_tightening |
gauge | 1 while the GC governor holds GOGC below its baseline | |
urnet_gc_cpu_fraction |
gauge | Share of CPU spent in GC over the last window; absent until a window completes | |
urnet_conntrack_used_ratio |
gauge |
nf_conntrack_count over nf_conntrack_max (Linux, when conntrack is loaded). Overflow drops packets silently |
|
urnet_tcp_time_wait |
gauge | TCP sockets in TIME_WAIT (Linux) | |
urnet_tcp_orphans |
gauge | Orphaned TCP sockets (Linux) | |
urnet_goroutines |
gauge | Goroutines | |
urnet_pool_latency_ms |
gauge | Message pool average latency | |
urnet_loop_restarts_total |
counter | loop |
Times a supervised background loop ended unexpectedly and was restarted |
urnet_loop_up |
gauge | loop |
1 while a supervised loop is running, 0 while it is stopped or backing off |