-
Notifications
You must be signed in to change notification settings - Fork 2
Operations and Verification
English | 简体中文
A healthy container, an open port, valid YAML, or an HTTP 101 proves only one
layer. Production acceptance requires evidence from configuration through a
real two-client desktop session. This page gives a repeatable order and clear
stop conditions.
| Level | What it proves | What it does not prove |
|---|---|---|
| Static | Compose and Nginx/config syntax is valid. | A process can start or a route is reachable. |
| Process | HBBS/HBBR are running and persistent files exist. | RustDesk protocol registration or Relay flow. |
| Network | DNS, TCP/UDP, and TLS endpoints are reachable. | Correct peer identity or message exchange. |
| Protocol | Clients register and a Relay is allocated. | Desktop data remains usable under load. |
| Session | Two real clients complete the intended path. | Failover and recovery. |
| Resilience | Failure and recovery behave as designed. | Future releases; repeat after every material change. |
Keep these outcomes separate in an acceptance record: passed, failed, and not tested. Never convert “not tested” into “passed”.
Before starting or changing services, record:
docker compose --env-file .env -f compose.yaml config --images
docker compose --env-file .env -f compose.yaml config > rendered-compose.review.txt
sha256sum .env compose.yaml data/starry/config.yaml > deployment-inputs.sha256
docker image inspect \
ghcr.io/q1ngyang/rustdesk-server-starry:1.1.16-patch-v1.2.0 \
--format '{{json .RepoDigests}}'Review the rendered Compose file before retaining it: it may include values
from .env. Store secrets and review evidence in protected operator storage,
not in the repository or an issue.
Also record the official upstream version, Starry patch version, architecture, public HBBS/HBBR names, and the expected client test matrix.
docker compose --env-file .env -f compose.yaml config --quiet
sudo nginx -tFor a centre/Relay topology, validate every file independently:
docker compose --env-file examples/center/.env \
-f examples/center/compose.bootstrap.yaml config --quiet
docker compose --env-file examples/center/.env \
-f examples/center/compose.yaml config --quiet
docker compose --env-file examples/relay/.env \
-f examples/relay/compose.yaml config --quietDo not start the full centre file until bootstrap has generated
id_ed25519.pub and the optional API mount can resolve it. A successful
Compose render does not check whether the bind-mounted public-key file is
already present.
docker compose --env-file .env -f compose.yaml ps
docker compose --env-file .env -f compose.yaml logs --tail 200 hbbs hbbr
test -s data/id_ed25519
test -s data/id_ed25519.pub
test -f data/starry/config.yaml
test -f data/starry/config.example.yamlExpected:
- HBBS and HBBR remain running rather than restarting;
- the HBBS health check becomes healthy;
- Starry logs show whether the external config was accepted;
- HBBR uses the same persistent public/private key pair as HBBS; and
- keys and database files survive a controlled container recreation.
Back up id_ed25519 securely. Distribute only id_ed25519.pub. A new private
key changes the server identity and breaks clients configured for the old key.
From a network outside the server's LAN, verify TCP reachability:
nc -vz id.example.com 21115
nc -vz id.example.com 21116
nc -vz relay.example.com 21117On Windows PowerShell:
Test-NetConnection id.example.com -Port 21116
Test-NetConnection relay.example.com -Port 21117UDP 21116 needs a registration/heartbeat observation or packet capture; a
generic TCP test cannot validate UDP. Confirm firewall and cloud security-group
rules in both directions and correlate the test time with HBBS logs.
Never expose 21118 or 21119 directly when Nginx is the intended public
entry point. Bind or firewall them to the reverse-proxy path.
Check DNS and the public certificate before the Upgrade request:
openssl s_client -connect id.example.com:443 \
-servername id.example.com -verify_return_error </dev/null
openssl s_client -connect relay.example.com:443 \
-servername relay.example.com -verify_return_error </dev/nullCheck the exact public paths using HTTP/1.1:
curl --http1.1 -i --max-time 5 \
-H 'Connection: Upgrade' \
-H 'Upgrade: websocket' \
-H 'Sec-WebSocket-Version: 13' \
-H 'Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==' \
https://id.example.com/ws/id
curl --http1.1 -i --max-time 5 \
-H 'Connection: Upgrade' \
-H 'Upgrade: websocket' \
-H 'Sec-WebSocket-Version: 13' \
-H 'Sec-WebSocket-Key: dGhlIHNhbXBsZSBub25jZQ==' \
https://relay.example.com/ws/relay101 Switching Protocols proves TLS termination and the Upgrade route. Curl
does not send a RustDesk registration or Relay handshake, so the expected later
timeout is not a real-session failure and the initial 101 is not full proof.
Inspect Starry's certificate-valid Relay probes through the authenticated
Control Agent GET /control/v1/status response.
Wait at least one configured probe interval after enabling or changing
endpoints. Every Relay intended for wss should be healthy; a mixed session
also requires that same Relay to be native-online.
Use authenticated Control Agent methods: POST /control/v1/runtime:reload,
GET /control/v1/relays, and POST /control/v1/allocations:simulate with the
two observed addresses, transport, expected generation, and explain: true.
Replace the reserved addresses with the two public addresses actually observed
by HBBS. Repeat allocation simulation with wss and mixed when those paths are enabled.
Record the expected rule and Relay before running the command.
Configure two clients with the same ID Server and public key. Leave Relay Server empty for HBBS allocation and disable WebSocket for this baseline.
Run and record:
- both clients show ready and appear in HBBS logs;
- direct P2P works when the network allows it;
- a forced-Relay test reaches the intended HBBR on
21117/TCP; - keyboard, pointer, clipboard, and the features relevant to your environment work for a sustained period; and
- reconnect works after closing the session.
If a third-party API is deployed, separately verify API login and then a native
session after login. Evidence that /api/status or login succeeds does not
prove HBBS Secure TCP. Correlate the native 21116/TCP handshake and client
session logs.
Enable WebSocket on clients only after the native baseline passes. Test each matrix row independently:
| Client A | Client B | Required Relay eligibility | Expected result |
|---|---|---|---|
| WSS | WSS | WSS healthy | Session uses HBBR and remains usable. |
| WSS | Native | Native-online and WSS healthy | Mixed Relay session succeeds. |
| Native | WSS | Native-online and WSS healthy | Reverse mixed direction also succeeds. |
| Native | Native | Native-online | Existing native behaviour remains unchanged. |
For each row, capture:
- both client log excerpts with timestamps;
- HBBS registration/routing log lines;
- the selected HBBR log lines for the same Relay session; and
- at least one usable desktop/control observation.
An HTTP 101 without client RegisterPk and a complete HBBR session is not a
pass.
Use a maintenance window and stop only a Relay that is authorised for testing. Do not test failover by disrupting an unrelated production node.
For each applicable transport:
- establish the baseline on the first ordered Relay;
- stop or firewall that Relay in a controlled, reversible way;
- wait for the official native state and/or configured WSS failure threshold;
- verify allocation simulation now selects the next ordered eligible Relay;
- establish a new real session and confirm it reaches the fallback;
- restore the first Relay;
- wait for the success threshold; and
- verify new sessions return to the first priority Relay.
Existing sessions may terminate during Relay loss; the acceptance target is a correctly allocated new session unless your own availability design promises more.
Follow the full matrix in
Connection Authentication.
Start with schema v3 mode: audit; capture status before and after each
native/Secure/WSS PunchHoleRequest and direct RequestRelay case. Confirm
that would-deny increments without changing the expected session outcome, and
that missing/invalid requests do not reach an existing target or allocate a
Relay.
Do not mark enforce ready until supported real clients cover P2P, Relay, WSS/WSS, both mixed directions, logout/revoke/disable/password-reset, key rotation, introspection failure/recovery, and UDP no-allocation. A clean unit or synthetic transport test is necessary release evidence but not a substitute for this deployment matrix.
Commission the Agent with write_enabled: false. Verify mTLS CA/SAN failures,
service-JWT audience/azp/scope/expiry failures, read endpoints, and repeated
side-effect-free simulation. After staging enables writes, exercise ETag
mismatch, duplicate and conflicting idempotency keys, plan expiry, an accepted
apply with matching runtime acknowledgement, rollback-as-new-revision, HBBS
outage during apply, and Agent restart recovery.
Any manual_intervention_required, disk/runtime drift, missing audit intent,
or apply response without matching generation/digests is a hard stop. Use the
Control Agent recovery runbook
before accepting more writes.
Use a table like this for each release or material configuration change:
| Check | Expected | Result | Evidence/time |
|---|---|---|---|
| Compose and Nginx syntax | Valid | Not tested | |
| HBBS/HBBR stable and keys persistent | Yes | Not tested | |
| Native registration | Both clients | Not tested | |
| Native P2P | When network allows | Not tested | |
| Native Relay | Expected HBBR | Not tested | |
| Authenticated Secure TCP | When API/login is used | Not tested | |
| WSS-to-WSS | Expected HBBR | Not tested | |
| WSS-to-native | Both directions | Not tested | |
| Geo decisions | Every matrix row | Not tested | |
| Native failover/recovery | Ordered result | Not tested | |
| WSS/mixed failover/recovery | Ordered result | Not tested | |
| Connection auth audit matrix | Every expected decision classified | Not tested | |
| Connection auth enforce canary | No bypass; legitimate matrix passes | Not tested | |
| JWT key/introspection failure and recovery | Fail closed, then recover without restart | Not tested | |
| Allocation simulation purity | No production state/counter change | Not tested | |
| Agent read-only mTLS/RBAC | Every allow/deny case | Not tested | |
| Config apply/rollback/outage recovery | Ack or explicit safe rollback/block | Not tested | |
| Backup restore rehearsal | Keys/config/state recover | Not tested |
Sanitise peer IDs, tokens, public addresses when sharing evidence. Do not place private keys, API credentials, or full access tokens in logs or issues.
- Documentation home
- Getting started
- Docker image usage
- Docker deployment
- Native deployment
- Multi-node deployment
- Reverse proxy and TLS
- Client configuration
- Account/API integration
- Configuration reference
- Connection authentication
- Control Agent
- GEO rules: basics
- GEO rules: advanced
- Operations and verification
- Troubleshooting
- Upgrade and rollback
- Architecture and build
- 文档主页
- 快速开始
- Docker 镜像使用
- Docker 部署
- 原生部署
- 中心与 Relay 多节点部署
- 反向代理与 TLS
- 客户端配置
- 账户与 API 服务接入
- 配置参数详解
- 连接认证
- Control Agent
- Geo 规则:入门
- Geo 规则:进阶
- 运维与完整验证
- 常见问题排查
- 版本升级与回滚
- 架构与构建