Repository navigation
Configuration
Important
As of v3.23.0-fix.31.0, systemd and native installs configure the provider
through the control socket, not by hand-editing override.conf.
URNETWORK_PROFILE, URNETWORK_RAMLOGS, GOMEMLIMIT and GOGC are now held
in the provider's runtime state and set with urnet-tools:
urnet-tools turbo v8 # profile
urnet-tools ramlogs on # ramlogs
urnet-tools set gogc 200 # any runtime tuning keyThe environment variables below still work and remain the configuration
surface for Docker, where they are passed to the container. On systemd,
prefer urnet-tools: it is the single source of truth that urnet-tools set
and urnet-tools status both read, and a hand-edited drop-in can silently
disagree with it. See Control Socket & Runtime Settings.
Quick jump:
- Authentication & Identity
- Performance & Hardware Profiles
- Proxy Feeds & Reaping
- Self-Healing & Resource Pressure
- Monitoring & Telemetry
| Variable | Default | Description |
|---|---|---|
BUILD |
stable |
Set to jwt for auth code login, or stable for email/password auth. |
USER_AUTH |
- | Your email. Required if BUILD=stable. Also used for self-healing in BUILD=jwt mode to refresh expired tokens. |
PASSWORD |
- | Your password. Required if BUILD=stable. Also used for self-healing in BUILD=jwt mode to refresh expired tokens. |
URNETWORK_AUTH_CODE |
- | First-run auth code for BUILD=jwt. Use this instead of passing the code as a trailing command argument. Ignored once a JWT exists in the volume. |
UR_API_URL |
https://api.bringyour.com |
Custom API URL, for operators running their own backend. Must be set together with UR_CONNECT_URL. Applied once at startup via provider choose_network and persisted to ~/.urnetwork/network.json, so it survives restarts if that directory is on a volume, same as the JWT. |
UR_CONNECT_URL |
wss://connect.bringyour.com |
Custom connect (WebSocket signaling) URL, for operators running their own backend. Must be set together with UR_API_URL. See UR_API_URL. |
URNETWORK_NODE_NAME |
hostname / redacted IP | Friendly label for dashboard identity and webhook alerts. |
HOST_HOSTNAME |
- | Pass the host server name into the container. Use -e HOST_HOSTNAME=$(hostname) with docker run or HOST_HOSTNAME=${HOSTNAME} in Compose. |
| Variable | Default | Description |
|---|---|---|
URNETWORK_PROFILE |
- | Advanced provider profile: auto, lowmem, eco, turbo-v4, or turbo-v8. For turbo, prefer TURBO. On systemd/native, set this with urnet-tools turbo/eco/lowmode/auto rather than a drop-in (v31+). |
TURBO |
- | Set to v4 or v8 to enable turbo mode. Prefer this variable for Docker turbo mode. |
URNETWORK_RAMLOGS |
0 |
Set to 1 to redirect provider logs to RAM instead of stdout. Cannot be used with Docker --log-opt. On systemd/native, set this with urnet-tools ramlogs on|off (v31+). |
URNETWORK_MESSAGE_POOL_SHARD_COUNT |
16 |
Number of internal mutex shards per message-pool size class. Higher values reduce lock contention at high packet rates. Must be a power of two, 1β256. Set to 1 to disable sharding (pre-v24.35 behavior). Sane values: 8 (moderate), 16 (default), 32 (high-pps tier3+). |
URNETWORK_METRICS |
- | Exact bind address for the Prometheus /metrics endpoint, for example 100.64.0.10:9100. Overrides urnet-tools metrics listen. Leave it unset and urnet-tools metrics on chooses: loopback plus the machine's Tailscale address on bare metal, every interface inside a container. Never bind a public interface. See Monitoring. |
URNETWORK_SKIP_AUDIT |
0 |
Set to 1 to skip the startup system audit (disk speed benchmark, ulimit, conntrack checks). Useful in Docker where host sysctls aren't visible. |
GOTRACEBACK |
- | Set to crash to produce full goroutine stack traces on Go runtime crashes. Add Environment="GOTRACEBACK=crash" to the systemd override.conf. |
| Variable | Default | Description |
|---|---|---|
PROXY_URL |
- | Live proxy list URL, fetched and merged on an interval. Comma-separate for multiple sources. See Proxy URL Sources. |
PROXY_URL_REFRESH |
1h |
How often PROXY_URL is re-fetched to add new entries. |
PROXY_URL_MAX |
500 |
Caps total proxies sourced from PROXY_URL. 0 = unlimited. |
PROXY_DEAD_CLEANUP_SCOPE |
url |
none, url, or all: which proxies the automatic dead-proxy cleanup may touch. Default url means only URL-sourced dead proxies are cleaned up automatically. |
PROXY_DEAD_CLEANUP_INTERVAL |
6h |
Base cadence of the automatic cleanup job, when scope is not none (shrinks under pressure when self-heal is on). The reaper runs this often to sweep dead proxies from the allowed sources. |
| Variable | Default | Description |
|---|---|---|
URNETWORK_SELF_HEAL |
0 (off) |
Set to 1 to enable the pressure-based self-heal system: proportional URL-fetch pacing, probe concurrency scaling, pressure-scaled cleanup/reaper cadence, and AIMD proxy-pool sizing. Off by default: with self-heal off, every actuator behaves exactly as it did before this system existed. Toggle at runtime with urnet-tools self-heal on, urnet-tools self-heal off, or urnet-tools self-heal status (no restart required; the monitor starts sensing within ~30s). |
URNETWORK_H3 |
off |
Beta (v3.23.0-fix.32.9). Set to on (or 1, true, yes) to run an H3 (QUIC) transport beside H1 for the direct identity only, the one that reaches the platform from the host's own address. A proxied identity never gets it. A node with no proxy source configured (direct-only) gets H3, h3_datagram and h3_datagram_send switched on automatically once the proxy reload settles, because that node has a single identity and the extra socket is the whole cost. Only keys you have not set are filled in, so an explicit urnet-tools set h3 off always wins, and a node with proxies is never touched. H1 stays the health signal: if H3 cannot connect (for example UDP is filtered) it backs off quietly up to 10 minutes and is not counted as a backend or proxy failure, and you see one [t]h3 unavailable, staying on h1 line. Costs one extra platform connection per box, and its socket bytes count into the identity's total traffic, not billable. No platform-side gain is proven. This is only the startup default: the h3 control key overrides it and switches H3 on or off live (see below). |
URNETWORK_PROXY_AUDIT |
0 (off) |
Set to 1 to enable automated proxy audit and quality enforcement: parks paid and file proxies that grade as proven junk. Off by default (observe mode). Toggle at runtime with urnet-tools proxy audit on|off|status|release without restarting or dropping sessions. |
URNETWORK_OOM_CAP |
shadow |
OOM-aware start cap. shadow (default) decides and logs what it would do after a kernel OOM kill since the previous start; on enforces the automatic cap (the tighter of it and your proxy trim cap applies); off disables it and forgets any standing automatic cap, so a cap set days earlier does not resume when you switch back on. Runtime switch: urnet-tools set oom-cap on|off|shadow; any source saying off wins. Decisions are recorded in ~/.urnetwork/autopilot.jsonl. See the capacity control log lines. |
baseline (control key) |
on |
Whether the provider records its own behaviour in ~/.urnetwork/baseline.jsonl. On by default, because a free upgrade baseline is the point: you should not have to configure anything to be able to judge the next upgrade. urnet-tools set baseline off stops recording at once, without a restart, and keeps the existing file: it is the only record of this box's behaviour before an upgrade. Clearing the key restores on. Read with urnet-tools baseline show and compared with urnet-tools baseline compare. |
URNETWORK_BASELINE_INTERVAL |
15m |
How often the baseline recorder samples itself, for a faster canary or a shorter soak. The first sample is 5 minutes after start rather than immediately, because a sample at t=0 measures a pool that has not launched yet. Values below one minute are clamped up to one minute. Changing it does not affect samples already recorded. |
URNETWORK_ADAPTIVE_GC |
on | Consolidated adaptive GC governor in the pressure monitor. On by default for every profile. It tightens GOGC below the profile baseline under memory pressure: the tighter of process heap fraction and host available RAM wins. Set to 0, false, off, or no to disable it. If the operator sets GOGC directly, the governor backs off entirely and never touches the knob. |
| Variable | Default | Description |
|---|---|---|
ENABLE_VNSTAT |
true |
Enables the traffic monitor on port 8080. |
ENABLE_IP_CHECKER |
false |
Diagnostic only. Prints your full public IP to container logs on startup via an external script. Distinct from dashboard identity reporting, which sends only a redacted IP. |
/metrics endpoint |
off |
Prometheus /metrics is off until it is turned on. Run urnet-tools metrics on to start the listener; it binds the first free port in 9100-9103, normally 9100. No URNETWORK_METRICS environment variable is required: the setting is persisted and re-applied at startup. The default bind is loopback, plus this machine's Tailscale address on bare metal and every interface inside a container, never a public one. Scrape http://<host>:9100/metrics from Prometheus or any compatible collector. If 9100 is already taken, the provider takes the next free port up to 9103, so read the address from urnet-tools metrics status instead of assuming 9100. Turn it off again with urnet-tools metrics off. URNETWORK_METRICS sets a listen address rather than a switch, so a value of 0 is an invalid address that fails to bind instead of a way to disable the endpoint. |
URNETWORK_PUBLIC_IP |
<auto-detected> |
Override the public IP shown in the dashboard identity label. Display only; does not change the actual egress IP. Auto-detected via ip.me on native/systemd installs; auto-set by Docker startup scripts. Create ~/.urnetwork/disable_ip_autodetect or run urnet-tools ip-detect off to prevent autodetection. See Node-Identity.md. |
URNETWORK_HEALTH_INTERVAL |
5m |
How often to emit a [health] heartbeat log line. Includes uptime, RAM stats, and active connection count. Accepts Go duration strings such as 10m or 1h. Minimum 1m. |
URNETWORK_PPROF |
- | Set to a host:port to enable the loopback-only diagnostics server (e.g. 127.0.0.1:6060). Off by default. Serves /debug/pprof/*, /metrics/pool, and /metrics/errors; only literal loopback IPs are accepted (hostnames are rejected). Pull profiles via an SSH tunnel, e.g. ssh -L 6060:127.0.0.1:6060 host then go tool pprof http://127.0.0.1:6060/debug/pprof/profile. |
URNETWORK_PROXY_BENCHMARK |
- | Set to true to enable per-proxy latency monitoring. Off by default. Probes: TCP connect every 5 min (raw RTT to proxy port), SOCKS5 CONNECT every 15 min (end-to-end through proxy). Staggered startup jitter prevents thundering herd. ~104 GB/month at 10k proxies. |
URNETWORK_PROXY_BENCHMARK_ENDPOINT |
connect.bringyour.com:443 |
Target for the SOCKS5 CONNECT latency probe. Measured end-to-end through each proxy. |
URNETWORK_REPORT_URL |
- | Startup fallback for the report URL. Precedence: control-socket state, then ~/.urnetwork/report_url, then this env var (read once at process start). Set per provider with urnet-tools report <url>; report off disables. |
URNETWORK_REPORT_INTERVAL |
5m |
Report cadence, minimum 10s. Re-resolved on every reporting tick, so runtime control-socket overrides apply without a restart. |
URNETWORK_HEARTBEAT_INTERVAL |
15s |
Heartbeat cadence (minimum 5s) for the lightweight heartbeat reporter that posts to the report target alongside the 5m full report. |
URNETWORK_AUTH_UNLIMITED |
false |
Bypass the auth rate limiter; every auth attempt fires immediately. Equivalent to creating ~/.urnetwork/fast_auth. Only for trusted or benchmark environments. |
URNETWORK_PUBLIC_IP |
<detected> |
Override the public IP shown in the dashboard identity label. Display only; does not change the actual egress IP. Auto-set by Docker startup scripts. |
URNETWORK_SHM_LOG |
/dev/shm/urnetwork.log |
Path for the RAM log. |
URNETWORK_PROXY_HEALTH_DIR |
<home>/.urnetwork |
Directory for persistent proxy_health.state and proxy_traffic.state files (Docker: /root/.urnetwork). |
URNETWORK_CONTAINER |
unset | Set to 1 to mark the process as containerized when /.dockerenv is absent (Podman, Kubernetes and other runtimes). A container counts as a restart supervisor for the thrash watchdog's exit 75. |
URNETWORK_CONTAINER_NAME |
<container-id> |
Container name used in copy-paste docker exec <name> tail -f hints for RAM logs. |
WARP_HOST |
<hostname> |
Override the host string reported by the warp status endpoint. Diagnostic only. |
Location: ~/.urnetwork/proxy_probe.json. Controls the stage-1 table-probe
quality gate for URL-source proxy admission (see
Proxy URL Sources). All
fields are optional; omitted fields keep their default. The file is
re-read on a short cache (a few seconds), so changes take effect without
a restart.
| Field | Type | Default | Meaning |
|---|---|---|---|
enabled |
bool | true |
Kill switch. false disables stage-1 entirely; proxies are admitted on stage-0 alone. |
sample_width |
int | 12 |
Intended sample base width, in destination-table hosts. |
min_sample_width |
int | 0 |
The small width a probe starts at. The paid grader forces 6. A clean verdict settles here and spends almost no probe bandwidth. |
max_sample_width |
int | 36 |
Upper bound adaptive sample growth may reach for a borderline proxy. |
timeout_ms |
int | 4000 |
Per-target dial timeout, in milliseconds. |
pass_bar |
float | 0.6 |
Minimum score (fraction of successful dials) required for admission to the auth queue. |
preferred_bar |
float | 0.9 |
Score threshold above which a proxy is marked preferred tier. |
borderline_band |
float | 0.15 |
Half-width around the pass bar that counts a proxy as borderline. A borderline score grows the sample toward max_sample_width; a score farther away is a decisive verdict and stops at the base width. |
use_spread_order |
bool | true |
Draw each sample block from a fixed, content-keyed spread of the destination table instead of consecutive rows. The table is grouped by theme, and a failure inside a theme is correlated, so a consecutive block could fail as a group and condemn a proxy whose capability never changed. false restores the consecutive-row sampler. |
min_confirm_dials |
int | 0 |
The fewest dials a below-bar pass must attempt before it may convict a proxy. The floor gates both the early abort and the final verdict, so a pass that runs out of block below the floor keeps its previous grade. The paid grader forces 6; the URL admission path and the reaper keep 0. -1 forces the floor off on the paid path. |
max_paid_probes_per_tick |
int | 200 |
Cap on how many paid/file proxies one 5-minute scoring sweep probes. |
stage0_liveness |
bool | false |
One-dial SOCKS5 and API reachability gate before a sample block. The paid grader forces true. |
Example, disabling the gate entirely:
{"enabled": false}Since v3.23.0-fix.25.14, the provider writes a per-process event log to ~/.urnetwork/events.log (on disk, not RAM β survives restarts). It records STARTUP, SIGNAL, PROVIDER EXIT, PANIC, and FATAL events. Capped at 1 MiB with automatic rotation.
cat ~/.urnetwork/events.log(v3.23.0-fix.31.0+, systemd and native installs)
| Path | Purpose |
|---|---|
~/.urnetwork/provider.sock |
Unix domain socket, owner-only 0600. The live control plane urnet-tools talks to. |
~/.urnetwork/provider_state.json |
Persisted runtime state. The provider is its single writer (atomic temp file + rename). |
~/.urnetwork/pending_overrides.json |
Queue for changes made while the provider is stopped. Flock-guarded, merged atomically on the next start, then removed. |
urnet-tools set # list current overrides
urnet-tools set report-interval 300 # change one, live, no restart
urnet-tools set report-interval off # clear itCapacity and dialer settings (v3.23.0-fix.32.8, v3.23.0-fix.32.9). Five settings are control keys. smart_dialer has its own command; the rest are set with urnet-tools set. All are live, persisted and re-applied at startup.
| Key | Values | Default | Command | What it does |
|---|---|---|---|---|
oom_cap |
on, off, shadow
|
shadow |
urnet-tools set oom-cap on|off|shadow (or URNETWORK_OOM_CAP in the unit) |
OOM-aware start cap. After a kernel OOM kill of the provider's own cgroup subtree the next start runs 80% of the peak running proxies, never below the larger of 50 and desired/4, at most 3 reductions per 24h, relaxing 10% at a start that comes a full clean day after the last change. shadow decides and logs and enforces nothing; on enforces; off disables it and forgets a standing cap. Any source saying off wins, even against URNETWORK_OOM_CAP=on. The effective cap is the tighter of this and your proxy trim cap. |
smart_dialer |
on, off
|
off |
urnet-tools smart-dialer [status|on|off] |
Transport choice from measured connect cost. With it on, weights come from success ratio and error streak, scaled by cost relative to the fastest transport that works here, and serial attempts try measured-faster transports first. It never removes a transport, a blocked transport still loses to a slow working one, and a client with one working transport is unaffected. There is no environment variable. With it off, scoring is exactly what it was. |
h3 |
on, off
|
follows URNETWORK_H3, else off; on automatically for a node with no proxy source configured (see below) |
urnet-tools set h3 on|off |
Beta. Switches the H3 (QUIC) transport on the direct identity on or off with no restart. On lets the idle transport dial; off closes a live H3 connection and stops further dials, and neither counts as an H3 drop. With it off, H1 behaves exactly as it does without H3. The key is persisted and re-applied at startup, so a persisted off beats URNETWORK_H3=on. urnet-tools set h3 off persists that off, while clearing the key hands the decision back to URNETWORK_H3. Its effect shows in [health] (h3_up, h3_tx_share, h3_drops, h3_conn_fail, appearing once H3 has been attempted) and in /metrics (urnet_transport_frames_total{mode,dir}, urnet_transport_payload_bytes_total{mode,dir}, urnet_h3_up, urnet_h3_connect_attempts_total, urnet_h3_connects_total, urnet_h3_connect_failures_total, urnet_h3_drops_total). |
h3_datagram |
on, off
|
off |
urnet-tools set h3-datagram on|off |
Experimental, receive side only. With it on, the H3 connection of the direct identity offers QUIC DATAGRAM (RFC 9221) when it dials. A server that accepts then sends its small frames to this provider as datagrams instead of on the reliable stream, so a lost packet no longer holds up the frames behind it. Everything this provider sends still goes on the stream unless h3_datagram_send is also on. A server that does not understand the offer echoes it back and the connection runs on the plain stream as before, which [health] shows as h3_dg=0/N. Changing it closes the live H3 connection and reconnects with the new setting, and that is not counted as a drop. It does nothing while h3 is off. There is no environment variable. Accepting changes how the server paces its sends to this provider, so enable it on one box first and watch [health] (h3_dg=accepted/offered, dg_rx, dg_rx_drop) and /metrics (urnet_h3_datagram_*). |
h3_datagram_send |
on, off
|
off |
urnet-tools set h3-datagram-send on|off |
Experimental. Lets this provider also SEND small frames as datagrams on an H3 connection where the server accepted DATAGRAM (h3_datagram); larger frames and everything else stay on the stream. It is read per message, so it takes effect at once with no reconnect. Datagrams are only used while H1 is up, because H1 is the authoritative path and the transfer layer resends whatever a lossy lane drops. A blackhole guard watches for datagrams going out and none coming back while the stream is alive, and when it fires the stream carries everything for the rest of that connection and dg_blackhole counts it. Stream-lane frames go through a queue bounded by count and by retained bytes. There is no environment variable. Watch [health] (dg_tx, dg_tx_stream, dg_tx_err, dg_blackhole) and /metrics (urnet_h3_datagram_tx_*, urnet_h3_datagram_blackholes_total). |
Note
The smart dialer is off by default in this release. The plan is to turn it on by default in a later release once it has had more testing, so enabling it on a few nodes now is what gets it there.
Read every capacity decision with urnet-tools autopilot log [limit]. The log lines these settings produce are described in the Log Message Reference and the smart dialer section.
Confirming a change landed. The provider logs every accepted and rejected change, which is the authoritative signal rather than the CLI's exit code:
βοΈ [control] set report-interval=300 (was unset)
βοΈ [control] applied 2 queued override(s) from pending_overrides.json: profile=v8, cleared gogc
β [control] set gogc=abc rejected: ...
profile and ramlogs still require the restart that urnet-tools performs
for you: buffer and worker sizing is baked into objects allocated once at
startup, and ramlogs is a live stdout redirect. The value is set through the
socket either way, so urnet-tools set and status stay the single source of
truth.
Tip
urnet-tools status reports whether the socket is actually bound. A running
PID with no reachable socket is a startup failure or a same-user collision,
not a healthy provider.
| Profile | Docker Value | Best For | RAM |
|---|---|---|---|
| Auto | URNETWORK_PROFILE=auto |
Recommended zero-config mode (auto-selects Low/Balanced/Perf/Extreme by RAM) | Any |
| Turbo V8 |
TURBO=v8 or URNETWORK_PROFILE=turbo-v8
|
Maximum throughput, dedicated servers | 16 GiB+ |
| Turbo V4 |
TURBO=v4 or URNETWORK_PROFILE=turbo-v4
|
High throughput, well-provisioned VPS | 4-16 GiB |
| Default | unset | General use | 2-4 GiB |
| Eco | URNETWORK_PROFILE=eco |
RAM-constrained, full throughput | 1-2 GiB |
| Lowmem | URNETWORK_PROFILE=lowmem |
Minimum RAM, reduced throughput | < 1 GiB |
See High-Volume Performance Tuning for the detailed profile behavior and parameter tables.
You can view the full list of dead and degraded proxies, as well as a live event log of proxy state transitions:
-
Host: Run
urnet-tools proxy health. -
Docker: See Docker Deployment for the
proxy-healthcommand.
Note
The proxy health files are stored in URNETWORK_PROXY_HEALTH_DIR (defaults to <home>/.urnetwork or /root/.urnetwork in Docker). Heartbeat intervals are tied to URNETWORK_HEALTH_INTERVAL (defaults to 5m).
Note
The status server (served on the provider's --port) sets ReadHeaderTimeout: 10s and IdleTimeout: 120s to prevent dribbled-header (Slowloris-style) clients from holding connections open indefinitely; WriteTimeout is deliberately unset so live streams are not killed.
URNETWORK_SELF_HEAL=1 (or urnet-tools self-heal on at runtime) turns on a resource-pressure monitor that scales several actuators proportionally instead of gating them on/off. It's off by default.
Sensors (sampled every 30s, worst-of-N combined, then smoothed with an asymmetric EWMA β fast to react, slow to relax):
-
/proc/pressure/memoryand/proc/pressure/cpu(PSIsome avg60), where available -
MemAvailable / MemTotalfrom/proc/meminfo -
loadavg1per core (fallback where PSI is unavailable) - Self-signals: goroutine count and heap fraction of the configured
max-memorysoft limit
These combine into a single smoothed pressure score in [0, 1]. A self-inflicted blowout (heap β₯90% of the soft limit, or β₯25,000 goroutines) pins the score to 1.0 immediately, bypassing smoothing.
Actuators, all driven off that one score:
- URL-fetch pacing stretches from 1Γ to 8Γ the configured interval as pressure rises (replaces the old binary skip-at-threshold gate)
- Proxy probe concurrency scales down toward a floor of 1 worker
- The dead-proxy cleanup job and the reaper's stale re-probe window both run more often under pressure (6h β 1h and 3h β 1h respectively) β cleanup and the reaper shed load, so pressure is exactly when they should run harder, not less
- An AIMD pool controller adjusts a persisted
TargetPoolSize(stored inproxy_url.json) every 5 minutes: +25 proxies when calm, Γ0.7 after two consecutive high-pressure samples (floor 50, capped byPROXY_URL_MAX). Shrinks evict the worst URL-sourced proxies first (dead, then degraded tiers, then healthy ones by ascending persisted earnings, then ascending lifetime traffic) with a 1h re-admission backoff. This learned target only caps admission while self-heal is enabled.
Everything above runs inside the provider. When the process is alive but too starved to run its own loops (a heap far over its limit on a small box leaves the garbage collector using most of the CPU), nothing in-process can rescue it, and systemd sees a running unit. The provider can hand that judgement to systemd: it sends WATCHDOG=1 only while its pressure monitor keeps ticking, and systemd restarts a unit that goes quiet. It is off until the unit sets WatchdogSec=, which needs the notify unit that urnet-tools update migrates to (Type=notify, NotifyAccess=all). A drop-in is enough:
# ~/.config/systemd/user/urnetwork.service.d/watchdog.conf (system units: /etc/systemd/system/urnetwork.service.d/)
[Service]
WatchdogSec=1200Then systemctl --user daemon-reload && systemctl --user restart urnetwork.service. The ping is withheld after 10 minutes without a tick, so systemd restarts a stalled provider between 10 and 30 minutes after the stall starts. The feed starts as soon as the provider reports ready, and until the pressure monitor has ticked once (it starts after the proxy list is loaded) startup gets a 30 minute grace, so a big list is not mistaken for a stall. That is deliberately slow: a long garbage collection pause or a briefly loaded box must never cost a restart. Before systemd acts, the provider records a lean start cap exactly as a thrash restart does (about 60% of the proxies that were running, counted in the same 3 per 24 hours ring), so the next start does not walk back into the same spiral. A cap already in force is kept, not cut again, so repeated stalls cannot compound down to one proxy, and the same ring and backoff as a thrash restart apply: once the day's restarts are spent, or inside the re-arm wait after one, the provider keeps feeding the watchdog and logs that the budget is spent, so a stall that recurs cannot restart the node for ever. The switch is the self-heal switch like everything else here: with self-heal off the provider keeps feeding the watchdog and only logs that it would have withheld the ping, because off means off for automatic restarts. If the stall resolves itself before systemd acts, the lean start cap it wrote is taken back out (and the restart ring entry with it). Keep WatchdogSec well above how long the provider takes to come up (the 20 minutes above does); the unit also needs NotifyAccess=main or all, which the notify unit has, and the provider says so in events.log if WatchdogSec= is set without it. Do not set WatchdogSec= on a unit whose binary predates this change: it never pings and systemd would restart it every interval.
Where there is no watchdog to starve the same detection still runs, with a different action: OpenRC's supervise-daemon and a container's start script both restart the provider on exit 75 but neither has an sd_notify watchdog, so after 10 minutes without a tick the provider exits with status 75 itself and its supervisor restarts it. It leaves the same lean start cap and counts against the same 3 per 24 hours ring, and the switch, the 30 minute startup grace and the resolution handling are identical to the systemd path; with self-heal off it logs what it would have done and keeps running. In a container the exit is refused when the state directory holding thrash_cap.json is not writable, because a mounted volume that does not survive the restart would reset the throttle on every restart.
A provider heading into a memory spiral stops answering within minutes: pprof times out, the log ring floods, the status files freeze, and by the time anyone looks the evidence of what piled up is gone. So the provider watches two cheap numbers every 10 seconds, its goroutine count and its heap against the soft limit, and when either shows the build-up it writes profiles to disk while it can still run:
| Trigger | When |
|---|---|
goroutine-growth |
at least 20,000 goroutines and 1.5 times the lowest count of the last 30 minutes (needs 10 minutes of history, and is not judged during the first 45 minutes after a start, while the proxies are still authenticating and the count climbs on its own) |
heap-near-limit |
heap at 90% of its soft limit |
heap-over-limit |
heap at 120% of its soft limit (heap profile and numbers only: the goroutine profile allocates a record per goroutine and is left out once the heap is past its limit) |
Each capture is a directory under ~/.urnetwork/incidents/<UTC time>-<trigger>/ with summary.txt (the 15 biggest goroutine stacks and how many goroutines share each, which is usually the whole answer to "what piled up"), goroutines.txt (the full grouped profile), heap.pb.gz (open it with go tool pprof) and meta.txt. A trigger repeats at most once per 30 minutes, captures are at least a minute apart, there are at most 6 a day, and the newest 8 directories are kept. It writes to disk only, never to the ramlog, and the profiles briefly pause the program while they are read (the heap profile and the numbers are cheap; the goroutine profile costs more the more goroutines there are). URNETWORK_INCIDENT_CAPTURE=0 turns it off. The line [incident] <trigger>: ... in ~/.urnetwork/events.log says when one was taken.
Memory pressure alone does not mean the box is thrashing. The watchdog watches for the signs that it is: PSI memory full (tasks stalled on memory), swap activity (pswpout/pswpin rates), page refaults, and direct reclaim. It keeps one state, calm -> under-pressure -> thrashing -> critical, and acts in steps:
- Freeze growth first: no new pool admissions while the condition holds. Nothing is lost yet.
- If the condition persists, the watchdog restarts the provider in a supervised way: the process exits with status 75 and the service manager restarts it. This is the only automatic restart the watchdog performs. It is capped at 3 restarts per 24 hours with a growing backoff (30m, 2h β a third rung is reserved for if the daily ceiling ever rises), skipped while a hot-swap is draining, and it refuses to act when the swap belongs to another process on the box.
- A restart leaves a thrash cap (
~/.urnetwork/thrash_cap.json): the next start begins with a smaller pool so it fits in RAM β about 60% of what was running before the restart, at least one proxy β and the cap expires 24 hours after the last restart. Nothing running before the restart sets no new cap (there is nothing to protect).
The supervised restart needs a service unit that restarts on exit status 75. The shipped units use Restart=on-failure, which covers it. If you override the unit with a drop-in, use Restart=on-failure or Restart=always, keep 75 (or its name TEMPFAIL) out of RestartPreventExitStatus, and do not mark 75 a success in SuccessExitStatus β under Restart=on-failure systemd would then treat the watchdog exit as a clean stop. The installer warns when a drop-in weakens any of this.
In a container the watchdog counts the container as a supervisor, the same as systemd and OpenRC, with no extra switch beyond self-heal. A container is recognised by /.dockerenv, or by URNETWORK_CONTAINER=1 on runtimes without that file (Podman, Kubernetes). Use a restart policy that restarts on exit (restart: unless-stopped or always, or the shipped start scripts, which restart on exit 75 after 5 seconds). Before exiting, the watchdog proves that ~/.urnetwork can be written and read back within a short timeout. If it cannot, it logs a persist-failed alert and does not restart, because a self-exit without its throttle record could loop. Keep ~/.urnetwork on a persistent volume so the 3 restarts per 24 hours ceiling survives the restart.
A heap running away. PSI can still read calm in the minutes before a spiral: the runtime pins a core on garbage collection, part of the heap swaps out, and soon the process is too starved to run any responder. So a heap at least 1.4 times its soft limit on a box with little free RAM (under a tenth of the box, never under 256 MiB) counts as severe on its own clock, 90 seconds, whatever PSI says. It needs both readings; a box that cannot say how much RAM is free is never judged on the heap alone.
All of this rides the existing self-heal switch (URNETWORK_SELF_HEAL=1 or urnet-tools self-heal on). Off means off for actions: with self-heal off, the watchdog still senses and logs, so you can watch it work, but it never restarts anything.
Where to look:
-
urnet-tools status's live block prints the current reading as one sentence (thesummaryfield the provider persists to~/.urnetwork/pressure_status). -
~/.urnetwork/thrash_statuscarries the machine-readable state for tooling. - Log lines start with
[proxy][thrash].
The paid and file proxies you supply are graded A to F by a probe from this box. With proxy audit on, the audit engine uses that grade to rest proxies that are proven junk, so their slots are not wasted. With proxy audit off it only watches and logs what it would have done, so you can see its judgement before you let it act. It also stays in observe mode when hot restart is off, because every relaunch would then mint a new client identity.
A proxy is parked (stopped, and not relaunched until its backoff ends) only when all of these hold:
- Its last two grades, from separate probe passes taken after the provider started, both scored 0.4 or lower. Tier F is anything under 0.6, but only 0.4 or lower counts as proven junk, and any grade above 0.4 in between resets the count.
- It is running, has no active clients, and has not earned recently. Idle alone never parks a proxy. A well-graded idle proxy is simply not being assigned traffic and is left alone; idle and earnings can only protect a proxy, never condemn it.
- The pass can be trusted: at least half of your proxies could be graded from this box, and there was no sudden mass failure (more than 40% of fresh grades bad at once, or twice the recent norm), which points at the box or its network rather than at the proxies.
Limits: at most 3 parks per 5-minute pass, at most 15% of your paid set per 24 hours (rounded up, so always at least 1), and proxy audit never parks a box down to fewer than max(10, half) running paid proxies. A box with 10 or fewer running paid proxies therefore never parks anything.
Backoff grows each time the same proxy is parked again: 6h, 12h, 24h, 48h, then 7 days, each varied by up to 25% so a batch does not return at once. A parked proxy returns early when it regrades to 0.6 or better after at least an hour. Turning proxy audit off returns everything it parked: it is released within one 5-minute pass, and the proxy is relaunched by the next reload, within about an hour (a reload is forced hourly). A regrade-based restore has the same lag. Parks are held in memory, so a provider restart returns them all and the evidence starts again.
A parked proxy is marked parked in proxy.state. It stays in your proxy file. Dead-proxy cleanup and proxy remove-dead never collect it in any category (dead, inactive, degraded or auth-failing, including the auth-failure check that --degraded turns on by default).
urnet-tools proxy audit status shows whether audit is observing or acting (and, when observing, whether proxy audit or hot restart is the switch that is off), what is parked and for how long, and why a pass parked nothing. On Linux, urnet-tools status shows the count in the unit's status line (3 parked by proxy audit); parked proxies do not make it read partial.
Proxy audit pauses when your paid proxy list cannot be read or is empty, since that list is what proves a proxy is yours to park. That is usually one skipped 5-minute pass while the file is being edited, but a missing or unreadable file keeps it paused until fixed. It logs [proxy][audit] paused and resumed, shows proxy audit: PAUSED in proxy audit status, and adds proxy audit paused to the unit's status line. Turning proxy audit off still releases parks while paused. The same is exported as urnet_proxy_audit_* gauges (see Monitoring). Log lines start with [proxy][audit].
Toggle or inspect at runtime:
urnet-tools proxy audit on
urnet-tools proxy audit off
urnet-tools proxy audit status
urnet-tools proxy audit release <addr>
urnet-tools proxy audit release --allNote
The ramp anchors (PSI 10%/60%, MemAvailable 25%/5%, load 1.0/3.0 per core, etc.) are properties of what each metric means β e.g. "a box stalled on memory 60% of the time is exhausted" holds regardless of core count or RAM size. They are not per-server capacity tuning knobs.
- Configuration Reference
- Node Identity
- HotSwap
- urnet-tools (Go)
- urnet-docker
- Proxy Management
- Proxy URL Sources
- Proxy Hot Reload
- Proxy Admission Pipeline
- High-Volume Tuning
- Bittensor Operations (Subnet 25)