# ⚙️ Configuration Reference > [!IMPORTANT] > **As of v3.23.0-fix.31.0, systemd and native installs configure the provider > through the control socket, not by hand-editing `override.conf`.** > `URNETWORK_PROFILE`, `URNETWORK_RAMLOGS`, `GOMEMLIMIT` and `GOGC` are now held > in the provider's runtime state and set with `urnet-tools`: > > ```bash > urnet-tools turbo v8 # profile > urnet-tools ramlogs on # ramlogs > urnet-tools set gogc 200 # any runtime tuning key > ``` > > The environment variables below still work and remain the configuration > surface for **Docker**, where they are passed to the container. On systemd, > prefer `urnet-tools`: it is the single source of truth that `urnet-tools set` > and `urnet-tools status` both read, and a hand-edited drop-in can silently > disagree with it. See [Control Socket & Runtime Settings](#-control-socket--runtime-settings). ## 🌍 Environment Variables Quick jump: - [Authentication & Identity](#-authentication--identity) - [Performance & Hardware Profiles](#-performance--hardware-profiles) - [Proxy Feeds & Reaping](#-proxy-feeds--reaping) - [Self-Healing & Resource Pressure](#-self-healing--resource-pressure) - [Monitoring & Telemetry](#-monitoring--telemetry) ### 🔑 Authentication & Identity | Variable | Default | Description | | :--- | :--- | :--- | | `BUILD` | `stable` | Set to `jwt` for auth code login, or `stable` for email/password auth. | | `USER_AUTH` | - | Your email. Required if `BUILD=stable`. Also used for **self-healing** in `BUILD=jwt` mode to refresh expired tokens. | | `PASSWORD` | - | Your password. Required if `BUILD=stable`. Also used for **self-healing** in `BUILD=jwt` mode to refresh expired tokens. | | `URNETWORK_AUTH_CODE` | - | First-run auth code for `BUILD=jwt`. Use this instead of passing the code as a trailing command argument. Ignored once a JWT exists in the volume. | | `UR_API_URL` | `https://api.bringyour.com` | Custom API URL, for operators running their own backend. Must be set together with `UR_CONNECT_URL`. Applied once at startup via `provider choose_network` and persisted to `~/.urnetwork/network.json`, so it survives restarts if that directory is on a volume, same as the JWT. | | `UR_CONNECT_URL` | `wss://connect.bringyour.com` | Custom connect (WebSocket signaling) URL, for operators running their own backend. Must be set together with `UR_API_URL`. See `UR_API_URL`. | | `URNETWORK_NODE_NAME` | hostname / redacted IP | Friendly label for dashboard identity and webhook alerts. | | `HOST_HOSTNAME` | - | Pass the host server name into the container. Use `-e HOST_HOSTNAME=$(hostname)` with `docker run` or `HOST_HOSTNAME=${HOSTNAME}` in Compose. | ### ⚡ Performance & Hardware Profiles | Variable | Default | Description | | :--- | :--- | :--- | | `URNETWORK_PROFILE` | - | Advanced provider profile: `auto`, `lowmem`, `eco`, `turbo-v4`, or `turbo-v8`. For turbo, prefer `TURBO`. On systemd/native, set this with `urnet-tools turbo/eco/lowmode/auto` rather than a drop-in (v31+). | | `TURBO` | - | Set to `v4` or `v8` to enable turbo mode. Prefer this variable for Docker turbo mode. | | `URNETWORK_RAMLOGS` | `0` | Set to `1` to redirect provider logs to RAM instead of stdout. Cannot be used with Docker `--log-opt`. On systemd/native, set this with `urnet-tools ramlogs on\|off` (v31+). | | `URNETWORK_MESSAGE_POOL_SHARD_COUNT` | `16` | Number of internal mutex shards per message-pool size class. Higher values reduce lock contention at high packet rates. Must be a power of two, 1–256. Set to `1` to disable sharding (pre-v24.35 behavior). Sane values: `8` (moderate), `16` (default), `32` (high-pps tier3+). | | `URNETWORK_METRICS` | - | Exact bind address for the Prometheus `/metrics` endpoint, for example `100.64.0.10:9100`. Overrides `urnet-tools metrics listen`. Leave it unset and `urnet-tools metrics on` chooses: loopback plus the machine's Tailscale address on bare metal, every interface inside a container. Never bind a public interface. See [Monitoring](Monitoring.md). | | `URNETWORK_SKIP_AUDIT` | `0` | Set to `1` to skip the startup system audit (disk speed benchmark, ulimit, conntrack checks). Useful in Docker where host sysctls aren't visible. | | `GOTRACEBACK` | - | Set to `crash` to produce full goroutine stack traces on Go runtime crashes. Add `Environment="GOTRACEBACK=crash"` to the systemd override.conf. | ### 🌐 Proxy Feeds & Reaping | Variable | Default | Description | | :--- | :--- | :--- | | `PROXY_URL` | - | Live proxy list URL, fetched and merged on an interval. Comma-separate for multiple sources. See [Proxy URL Sources](Proxy-URL-Sources.md). | | `PROXY_URL_REFRESH` | `1h` | How often `PROXY_URL` is re-fetched to add new entries. | | `PROXY_URL_MAX` | `500` | Caps total proxies sourced from `PROXY_URL`. `0` = unlimited. | | `PROXY_DEAD_CLEANUP_SCOPE` | `url` | `none`, `url`, or `all`: which proxies the automatic dead-proxy cleanup may touch. Default `url` means only URL-sourced dead proxies are cleaned up automatically. | | `PROXY_DEAD_CLEANUP_INTERVAL` | `6h` | Base cadence of the automatic cleanup job, when scope is not `none` (shrinks under pressure when self-heal is on). The reaper runs this often to sweep dead proxies from the allowed sources. | ### 🩹 Self-Healing & Resource Pressure | Variable | Default | Description | | :--- | :--- | :--- | | `URNETWORK_SELF_HEAL` | `0` (off) | Set to `1` to enable the pressure-based self-heal system: proportional URL-fetch pacing, probe concurrency scaling, pressure-scaled cleanup/reaper cadence, and AIMD proxy-pool sizing. Off by default: with self-heal off, every actuator behaves exactly as it did before this system existed. Toggle at runtime with `urnet-tools self-heal on`, `urnet-tools self-heal off`, or `urnet-tools self-heal status` (no restart required; the monitor starts sensing within ~30s). | | `URNETWORK_H3` | `off` | Beta (v3.23.0-fix.32.9). Set to `on` (or `1`, `true`, `yes`) to run an H3 (QUIC) transport beside H1 for the **direct** identity only, the one that reaches the platform from the host's own address. A proxied identity never gets it. H1 stays the health signal: if H3 cannot connect (for example UDP is filtered) it backs off quietly up to 10 minutes and is not counted as a backend or proxy failure, and you see one `[t]h3 unavailable, staying on h1` line. Costs one extra platform connection per box, and its socket bytes count into the identity's total traffic, not billable. No platform-side gain is proven. This is only the startup default: the `h3` control key overrides it and switches H3 on or off live (see below). | | `URNETWORK_PROXY_AUDIT` | `0` (off) | Set to `1` to enable automated [proxy audit](#proxy-audit) and quality enforcement: parks paid and file proxies that grade as proven junk. Off by default (observe mode). Toggle at runtime with `urnet-tools proxy audit on\|off\|status\|release` without restarting or dropping sessions. | | `URNETWORK_OOM_CAP` | `shadow` | OOM-aware start cap. `shadow` (default) decides and logs what it would do after a kernel OOM kill since the previous start; `on` enforces the automatic cap (the tighter of it and your `proxy trim` cap applies); `off` disables it and forgets any standing automatic cap, so a cap set days earlier does not resume when you switch back on. Runtime switch: `urnet-tools set oom-cap on\|off\|shadow`; any source saying `off` wins. Decisions are recorded in `~/.urnetwork/autopilot.jsonl`. See the [capacity control log lines](../LOG_REFERENCE.md). | | `URNETWORK_ADAPTIVE_GC` | on | Consolidated adaptive GC governor in the pressure monitor. On by default for every profile. It tightens GOGC below the profile baseline under memory pressure: the tighter of process heap fraction and host available RAM wins. Set to `0`, `false`, `off`, or `no` to disable it. If the operator sets `GOGC` directly, the governor backs off entirely and never touches the knob. | | `baseline` (control key) | `on` | Whether the provider records its own behaviour in `~/.urnetwork/baseline.jsonl`. On by default, because a free upgrade baseline is the point: you should not have to configure anything to be able to judge the next upgrade. `urnet-tools set baseline off` stops recording at once, without a restart, and keeps the existing file: it is the only record of this box's behaviour before an upgrade. Clearing the key restores `on`. Read with `urnet-tools baseline show` and compared with `urnet-tools baseline compare`. | | `URNETWORK_BASELINE_INTERVAL` | `15m` | How often the baseline recorder samples itself, for a faster canary or a shorter soak. The first sample is 5 minutes after start rather than immediately, because a sample at t=0 measures a pool that has not launched yet. Values below one minute are clamped up to one minute. Changing it does not affect samples already recorded. | ### 📊 Monitoring & Telemetry | Variable | Default | Description | | :--- | :--- | :--- | | `ENABLE_VNSTAT` | `true` | Enables the traffic monitor on port 8080. | | `ENABLE_IP_CHECKER` | `false` | Diagnostic only. Prints your full public IP to container logs on startup via an external script. Distinct from dashboard identity reporting, which sends only a redacted IP. | | `/metrics` endpoint | `:9091` | **Since v3.23.0-fix.31.2**, the Prometheus metrics endpoint is enabled by default on port `9091`. No `URNETWORK_METRICS` environment variable is required — the provider starts the metrics listener automatically. Scrape `http://:9091/metrics` from Prometheus or any compatible collector. To disable it, set `URNETWORK_METRICS=0`. | | `URNETWORK_PUBLIC_IP` | `` | Override the public IP shown in the dashboard identity label. Display only; does not change the actual egress IP. Auto-detected via `ip.me` on native/systemd installs; auto-set by Docker startup scripts. Create `~/.urnetwork/disable_ip_autodetect` or run `urnet-tools ip-detect off` to prevent autodetection. See [Node-Identity.md](Node-Identity.md). | | `URNETWORK_HEALTH_INTERVAL` | `5m` | How often to emit a `[health]` heartbeat log line. Includes uptime, RAM stats, and active connection count. Accepts Go duration strings such as `10m` or `1h`. Minimum `1m`. | | `URNETWORK_PPROF` | - | Set to a `host:port` to enable the loopback-only diagnostics server (e.g. `127.0.0.1:6060`). Off by default. Serves `/debug/pprof/*`, `/metrics/pool`, and `/metrics/errors`; only literal loopback IPs are accepted (hostnames are rejected). Pull profiles via an SSH tunnel, e.g. `ssh -L 6060:127.0.0.1:6060 host` then `go tool pprof http://127.0.0.1:6060/debug/pprof/profile`. | | `URNETWORK_PROXY_BENCHMARK` | - | Set to `true` to enable per-proxy latency monitoring. Off by default. Probes: TCP connect every 5 min (raw RTT to proxy port), SOCKS5 CONNECT every 15 min (end-to-end through proxy). Staggered startup jitter prevents thundering herd. ~104 GB/month at 10k proxies. | | `URNETWORK_PROXY_BENCHMARK_ENDPOINT` | `connect.bringyour.com:443` | Target for the SOCKS5 CONNECT latency probe. Measured end-to-end through each proxy. | | `URNETWORK_REPORT_URL` | - | *(Deprecated v31.3+)* HTTP URL of a bandwidth hub server. Was used to POST JSON reports with per-proxy metrics. See `docs/Hub-Dashboard.md` for historical reference. | | `URNETWORK_REPORT_INTERVAL` | `5m` | *(Deprecated v31.3+)* How often bandwidth reports were posted to `URNETWORK_REPORT_URL`. No longer functional. | | `URNETWORK_HEARTBEAT_INTERVAL` | `15s` | *(Deprecated v31.3+)* Provider heartbeat cadence to the hub. No longer functional. | | `URNETWORK_AUTH_UNLIMITED` | `false` | Bypass the auth rate limiter; every auth attempt fires immediately. Equivalent to creating `~/.urnetwork/fast_auth`. Only for trusted or benchmark environments. | | `URNETWORK_PUBLIC_IP` | `` | Override the public IP shown in the dashboard identity label. Display only; does not change the actual egress IP. Auto-set by Docker startup scripts. | | `URNETWORK_SHM_LOG` | `/dev/shm/urnetwork.log` | Path for the RAM log. | | `URNETWORK_PROXY_HEALTH_DIR` | `/.urnetwork` | Directory for persistent `proxy_health.state` and `proxy_traffic.state` files (Docker: `/root/.urnetwork`). | | `URNETWORK_CONTAINER_NAME` | `` | Container name used in copy-paste `docker exec tail -f` hints for RAM logs. | | `WARP_HOST` | `` | Override the host string reported by the warp status endpoint. Diagnostic only. | ## 📄 `proxy_probe.json` (Stage-1 Gate) Location: `~/.urnetwork/proxy_probe.json`. Controls the stage-1 table-probe quality gate for URL-source proxy admission (see [Proxy URL Sources](Proxy-URL-Sources.md#-stage-1-quality-gate)). All fields are optional; omitted fields keep their default. The file is re-read on a short cache (a few seconds), so changes take effect without a restart. | Field | Type | Default | Meaning | |---|---|---|---| | `enabled` | bool | `true` | Kill switch. `false` disables stage-1 entirely; proxies are admitted on stage-0 alone. | | `sample_width` | int | `12` | Intended sample base width, in destination-table hosts. | | `min_sample_width` | int | `0` | The small width a probe starts at. The paid grader forces 6. A clean verdict settles here and spends almost no probe bandwidth. | | `max_sample_width` | int | `36` | Upper bound adaptive sample growth may reach for a borderline proxy. | | `timeout_ms` | int | `4000` | Per-target dial timeout, in milliseconds. | | `pass_bar` | float | `0.6` | Minimum score (fraction of successful dials) required for admission to the auth queue. | | `preferred_bar` | float | `0.9` | Score threshold above which a proxy is marked preferred tier. | | `border_line_band` | float | `0.15` | Half-width around the pass bar that counts a proxy as borderline. A borderline score grows the sample toward `max_sample_width`; a score farther away is a decisive verdict and stops at the base width. | | `max_paid_probes_per_tick` | int | `200` | Cap on how many paid/file proxies one 5-minute scoring sweep probes. | | `stage0_liveness` | bool | `false` | One-dial SOCKS5 and API reachability gate before a sample block. The paid grader forces true. | Example, disabling the gate entirely: ```json {"enabled": false} ``` ## 📝 Critical Event Log Since v3.23.0-fix.25.14, the provider writes a per-process event log to `~/.urnetwork/events.log` (on disk, not RAM — survives restarts). It records STARTUP, SIGNAL, PROVIDER EXIT, PANIC, and FATAL events. Capped at 1 MiB with automatic rotation. ```bash cat ~/.urnetwork/events.log ``` ## 🔌 Control Socket & Runtime Settings *(v3.23.0-fix.31.0+, systemd and native installs)* | Path | Purpose | | :--- | :--- | | `~/.urnetwork/provider.sock` | Unix domain socket, owner-only `0600`. The live control plane `urnet-tools` talks to. | | `~/.urnetwork/provider_state.json` | Persisted runtime state. The provider is its single writer (atomic temp file + rename). | | `~/.urnetwork/pending_overrides.json` | Queue for changes made while the provider is stopped. Flock-guarded, merged atomically on the next start, then removed. | ```bash urnet-tools set # list current overrides urnet-tools set report-interval 300 # change one, live, no restart urnet-tools set report-interval off # clear it ``` **Capacity and dialer settings (v3.23.0-fix.32.8).** Two settings are control keys with their own commands. Both are live, persisted and re-applied at startup. | Key | Values | Default | Command | What it does | | :--- | :--- | :--- | :--- | :--- | | `oom_cap` | `on`, `off`, `shadow` | `shadow` | `urnet-tools set oom-cap on\|off\|shadow` (or `URNETWORK_OOM_CAP` in the unit) | OOM-aware start cap. After a kernel OOM kill of the provider's own cgroup subtree the next start runs 80% of the peak running proxies, never below the larger of 50 and desired/4, at most 3 reductions per 24h, relaxing 10% at a start that comes a full clean day after the last change. `shadow` decides and logs and enforces nothing; `on` enforces; `off` disables it and forgets a standing cap. Any source saying `off` wins, even against `URNETWORK_OOM_CAP=on`. The effective cap is the tighter of this and your `proxy trim` cap. | | `smart_dialer` | `on`, `off` | `off` | `urnet-tools smart-dialer [status\|on\|off]` | Transport choice from measured connect cost. With it on, weights come from success ratio and error streak, scaled by cost relative to the fastest transport that works here, and serial attempts try measured-faster transports first. It never removes a transport, a blocked transport still loses to a slow working one, and a client with one working transport is unaffected. There is no environment variable. With it off, scoring is exactly what it was. | | `h3` | `on`, `off` | follows `URNETWORK_H3`, else `off` | `urnet-tools set h3 on\|off` | Beta. Switches the H3 (QUIC) transport on the **direct** identity on or off with no restart. On lets the idle transport dial; off closes a live H3 connection and stops further dials, and neither counts as an H3 drop. With it off, H1 behaves exactly as it does without H3. The key is persisted and re-applied at startup, where a persisted `off` beats `URNETWORK_H3=on`; clearing it hands the decision back to `URNETWORK_H3`. Its effect shows in `[health]` (`h3_up`, `h3_tx_share`, `h3_drops`, `h3_conn_fail`, appearing once H3 has been attempted) and in `/metrics` (`urnet_transport_frames_total{mode,dir}`, `urnet_transport_payload_bytes_total{mode,dir}`, `urnet_h3_up`, `urnet_h3_connect_attempts_total`, `urnet_h3_connects_total`, `urnet_h3_connect_failures_total`, `urnet_h3_drops_total`). | | `h3_datagram` | `on`, `off` | `off` | `urnet-tools set h3-datagram on\|off` | Experimental, receive side only. With it on, the H3 connection of the direct identity offers QUIC DATAGRAM (RFC 9221) when it dials. A server that accepts then sends its small frames to this provider as datagrams instead of on the reliable stream, so a lost packet no longer holds up the frames behind it. Everything this provider sends still goes on the stream unless `h3_datagram_send` is also on. A server that does not understand the offer echoes it back and the connection runs on the plain stream as before, which `[health]` shows as `h3_dg=0/N`. Changing it closes the live H3 connection and reconnects with the new setting, and that is not counted as a drop. It does nothing while `h3` is off. There is no environment variable. Accepting changes how the server paces its sends to this provider, so enable it on one box first and watch `[health]` (`h3_dg=accepted/offered`, `dg_rx`, `dg_rx_drop`) and `/metrics` (`urnet_h3_datagram_*`). | | `h3_datagram_send` | `on`, `off` | `off` | `urnet-tools set h3-datagram-send on\|off` | Experimental. Lets this provider also SEND small frames as datagrams on an H3 connection where the server accepted DATAGRAM (`h3_datagram`); larger frames and everything else stay on the stream. It is read per message, so it takes effect at once with no reconnect. Datagrams are only used while H1 is up, because H1 is the authoritative path and the transfer layer resends whatever a lossy lane drops. A blackhole guard watches for datagrams going out and none coming back while the stream is alive, and when it fires the stream carries everything for the rest of that connection and `dg_blackhole` counts it. Stream-lane frames go through a queue bounded by count and by retained bytes. There is no environment variable. Watch `[health]` (`dg_tx`, `dg_tx_stream`, `dg_tx_err`, `dg_blackhole`) and `/metrics` (`urnet_h3_datagram_tx_*`, `urnet_h3_datagram_blackholes_total`). | > [!NOTE] > The smart dialer is off by default in this release. The plan is to turn it on by default in a later release once it has had more testing, so enabling it on a few nodes now is what gets it there. Read every capacity decision with `urnet-tools autopilot log [limit]`. The log lines these settings produce are described in the [Log Message Reference](../LOG_REFERENCE.md) and the [smart dialer section](../LOG_REFERENCE.md#-smart-dialer-and-give-up-lines). **Confirming a change landed.** The provider logs every accepted and rejected change, which is the authoritative signal rather than the CLI's exit code: ```text ⚙️ [control] set report-interval=300 (was unset) ⚙️ [control] applied 2 queued override(s) from pending_overrides.json: profile=v8, cleared gogc ❌ [control] set gogc=abc rejected: ... ``` `profile` and `ramlogs` still require the restart that `urnet-tools` performs for you: buffer and worker sizing is baked into objects allocated once at startup, and ramlogs is a live stdout redirect. The value is set through the socket either way, so `urnet-tools set` and `status` stay the single source of truth. > [!TIP] > `urnet-tools status` reports whether the socket is actually bound. A running > PID with no reachable socket is a startup failure or a same-user collision, > not a healthy provider. ## 🎛️ Profile Selection | Profile | Docker Value | Best For | RAM | | :--- | :--- | :--- | :--- | | Auto | `URNETWORK_PROFILE=auto` | Recommended zero-config mode (auto-selects Low/Balanced/Perf/Extreme by RAM) | Any | | Turbo V8 | `TURBO=v8` or `URNETWORK_PROFILE=turbo-v8` | Maximum throughput, dedicated servers | 16 GiB+ | | Turbo V4 | `TURBO=v4` or `URNETWORK_PROFILE=turbo-v4` | High throughput, well-provisioned VPS | 4-16 GiB | | Default | unset | General use | 2-4 GiB | | Eco | `URNETWORK_PROFILE=eco` | RAM-constrained, full throughput | 1-2 GiB | | Lowmem | `URNETWORK_PROFILE=lowmem` | Minimum RAM, reduced throughput | < 1 GiB | See [High-Volume Performance Tuning](High-Volume-Performance-Tuning.md) for the detailed profile behavior and parameter tables. ## 🩺 Viewing proxy health You can view the full list of dead and degraded proxies, as well as a live event log of proxy state transitions: * **Host**: Run `urnet-tools proxy health`. * **Docker**: See [Docker Deployment](Docker-Deployment.md) for the `proxy-health` command. > [!NOTE] > The proxy health files are stored in `URNETWORK_PROXY_HEALTH_DIR` (defaults to `/.urnetwork` or `/root/.urnetwork` in Docker). Heartbeat intervals are tied to `URNETWORK_HEALTH_INTERVAL` (defaults to 5m). > [!NOTE] > The status server (served on the provider's `--port`) sets `ReadHeaderTimeout: 10s` and `IdleTimeout: 120s` to prevent dribbled-header (Slowloris-style) clients from holding connections open indefinitely; `WriteTimeout` is deliberately unset so live streams are not killed. ## 🩹 Pressure system (self-heal) `URNETWORK_SELF_HEAL=1` (or `urnet-tools self-heal on` at runtime) turns on a resource-pressure monitor that scales several actuators proportionally instead of gating them on/off. It's off by default. **Sensors** (sampled every 30s, worst-of-N combined, then smoothed with an asymmetric EWMA — fast to react, slow to relax): - `/proc/pressure/memory` and `/proc/pressure/cpu` (PSI `some avg60`), where available - `MemAvailable / MemTotal` from `/proc/meminfo` - `loadavg1` per core (fallback where PSI is unavailable) - Self-signals: goroutine count and heap fraction of the configured `max-memory` soft limit These combine into a single smoothed pressure score in `[0, 1]`. A self-inflicted blowout (heap ≥90% of the soft limit, or ≥25,000 goroutines) pins the score to `1.0` immediately, bypassing smoothing. **Actuators**, all driven off that one score: - URL-fetch pacing stretches from 1× to 8× the configured interval as pressure rises (replaces the old binary skip-at-threshold gate) - Proxy probe concurrency scales down toward a floor of 1 worker - The dead-proxy cleanup job and the reaper's stale re-probe window both run *more* often under pressure (6h → 1h and 3h → 1h respectively) — cleanup and the reaper shed load, so pressure is exactly when they should run harder, not less - An AIMD pool controller adjusts a persisted `TargetPoolSize` (stored in `proxy_url.json`) every 5 minutes: +25 proxies when calm, ×0.7 after two consecutive high-pressure samples (floor 50, capped by `PROXY_URL_MAX`). Shrinks evict the worst URL-sourced proxies first (dead, then degraded tiers, then healthy ones by ascending traffic) with a 1h re-admission backoff. This learned target only caps admission while self-heal is enabled. Check current state with `urnet-tools self-heal status`, which prints the on/off toggle plus the live score, per-component breakdown, and target pool size from `~/.urnetwork/pressure_status`. The status file also reports `gc_state` and `heap_frac`, the adaptive GC governor's current level and live heap fraction. > [!NOTE] > The ramp anchors (PSI 10%/60%, MemAvailable 25%/5%, load 1.0/3.0 per core, etc.) are properties of what each metric means — e.g. "a box stalled on memory 60% of the time is exhausted" holds regardless of core count or RAM size. They are not per-server capacity tuning knobs. ### Swap-thrash watchdog Memory pressure alone does not mean the box is thrashing. The watchdog watches for the signs that it is: PSI memory `full` (tasks stalled on memory), swap activity (`pswpout`/`pswpin` rates), page refaults, and direct reclaim. It keeps one state, `calm -> under-pressure -> thrashing -> critical`, and acts in steps: - **Freeze growth** first: no new pool admissions while the condition holds. Nothing is lost yet. - If the condition persists, the watchdog **restarts the provider** in a supervised way: the process exits with status 75 and the service manager restarts it. This is the only automatic restart the watchdog performs. It is capped at 3 restarts per 24 hours with a growing backoff (30m, 2h — a third rung is reserved for if the daily ceiling ever rises), skipped while a hot-swap is draining, and it refuses to act when the swap belongs to another process on the box. - A restart leaves a **thrash cap** (`~/.urnetwork/thrash_cap.json`): the next start begins with a smaller pool so it fits in RAM — about 60% of what was running before the restart, at least one proxy — and the cap expires 24 hours after the last restart. Nothing running before the restart sets no new cap (there is nothing to protect). The supervised restart needs a service unit that restarts on exit status 75. The shipped units use `Restart=on-failure`, which covers it. If you override the unit with a drop-in, use `Restart=on-failure` or `Restart=always`, keep 75 (or its name `TEMPFAIL`) out of `RestartPreventExitStatus`, and do not mark 75 a success in `SuccessExitStatus` — under `Restart=on-failure` systemd would then treat the watchdog exit as a clean stop. The installer warns when a drop-in weakens any of this. All of this rides the existing self-heal switch (`URNETWORK_SELF_HEAL=1` or `urnet-tools self-heal on`). Off means off for actions: with self-heal off, the watchdog still senses and logs, so you can watch it work, but it never restarts anything. Where to look: - `urnet-tools status`'s live block prints the current reading as one sentence (the `summary` field the provider persists to `~/.urnetwork/pressure_status`). - `~/.urnetwork/thrash_status` carries the machine-readable state for tooling. - Log lines start with `[proxy][thrash]`.