Repository navigation
ripcord 0.8.0
What is in 0.8.0
Added
- A second service,
ripcord-publish, republishes this host's snapshot every 15 seconds.
The other host's view no longer ages until somebody runsripcord statushere. The
network-facing listener still reads nothing of Hyper-V; the publisher, which has no socket,
runs asNT SERVICE\ripcord-publishin Hyper-V Administrators, with modify onstate\and
on its ownlogs\publish\only, and one ACE on the BitLocker namespace while
storage.check_bitlocker_autounlockis on. BitLocker it still cannot read is carried from
the last administrator's read, dated.ripcord publishdoes it once by hand. service installandservice removehandle both services;restart,startand
stopact on both, andripcord serviceshows the publisher with the last lines of its log.
Install also restarts a service still running the build from beforeripcord update, sets
both to restart a minute after a crash, and narrows the listener's 0.7.0 read on the install
folder to its own files.
Upgrading a host: ripcord update, then ripcord service install --dry-run and
ripcord service install. It creates the publisher and state\, restarts the listener onto
the new binary and state\state.json, and starts the publisher. If a stale snapshot remains,
ripcord service names the one command that fixes it. The old state.json beside
ripcord.yaml can be deleted.
Changed
- The snapshot moves to
state\state.jsonbesideripcord.yaml, and is no longer
configurable. A folder of its own, so nothing that writes it is ever granted anything
besideripcord.exe. Alistener.snapshot_pathline still loads, is ignored, and every
command says so. See Upgrading a host above. service installgrants the listener read access toripcord.yamlexplicitly, on the
files of the install folder only. It used to come from the snapshot grant on that folder.
Fixed
-
The peer's lag grew with its snapshot's age. A snapshot 7 minutes old showed 7m19s of
lag on a replication 17 s behind: the lag was measured to now instead of to when the
snapshot was taken.statusand the dashboard now measure it at the snapshot. -
ripcord service installtakes back access nothing uses any more. The service account
kept read access to the private key of every certificate replaced bypairor a renewal,
and to the old install folder after the binary moved. Install now revokes them, last, after
everything the listener needs; with the configured key not found it touches no key.
Before you update
Update both hosts, one after the other, and finish. Ripcord's failover sequences are encoded
in the binary and they span two hosts, so a pair running two versions would execute half a
sequence written by each. A planned failover and a failback therefore refuse on a version
mismatch rather than warning — which means a pair left half-updated cannot be moved electively.
The window between the two hosts is the one time the tool will not work.
The one exception is failover --scenario unplanned, which runs entirely on the host it is
typed on. There is no half for the other binary to execute, and refusing a disaster failover for
want of a version string the dead host never published would fail at the one thing this tool
exists for.
Do it when nothing is on fire, not during an incident.
Verify what you are installing
There is no Authenticode signature yet; it needs a paid certificate. Until then the SHA-256
checksum published beside the binary is the only way to tell that the file on the host is the
file that was built. Check it before replacing anything.
(Get-FileHash ripcord.exe -Algorithm SHA256).Hash.ToLower()Updating
After updating a host, run ripcord service install --dry-run, then
ripcord service install. A release can need something new on the host; install applies it,
and restarts a Ripcord service still running the build from before the update. Then the other
host, the same way.
ripcord update installs a newer release on the host it is run on. It is off unless
updates.install says so, it asks y/n with Enter declining, and it refuses any release
whose detached signature does not verify against the key compiled into the running binary.
Copying the .exe by hand still works and is still supported; fully offline operation stays
possible.
This reverses what earlier versions of this file said. The objection was sound and has not gone
away: a binary that replaces itself, with Hyper-V privileges, on both hosts of a disaster
recovery pair, from the internet, is a supply chain reaching past the controls the rest of the
tool is built around. What changed is that the path now has a trust root that is not the release
page — the signature is checked against a pinned key, and an unverifiable release is refused
before anything is touched. What has not changed is that the signing key lives in the
release workflow, so an account with write access to the repository can still sign; see
SECURITY.md.
Update one host at a time, and update both. While the two differ, a failover spanning them
is refused — the sequences are encoded in the binary and one executed half by each version is
the error nobody recovers from at 3 a.m. The command says which host to run next.
The full procedure is in docs/RELEASING.md.