Repository navigation
ripcord 0.7.0
What is in 0.7.0
Changed
- Two confirmation levels. Only what moves production VMs —
failover,failback,
fence— still asks for the node name typed in full.service install,removeandstop,
test-failover,updateandrollbackasky/n [n]instead, Enter declining: each is
undone by running it again or touches only a test VM, and a typed name asked for every
change becomes the reflex the failover confirmation must never be.
Added
-
ripcord serviceshows this host's certificate: the thumbprint of each unexpired
certificate whose CN is this host's, with a private key inLocalMachine\My, marked when it
is the one inripcord.yaml, even when that file does not load — with the one line to run
on the other host. -
ripcord pair <host>:<thumbprint>writes both certificate thumbprints into
ripcord.yamlfrom that line: this host's own is found inLocalMachine\My, the other's is
the one pasted. Two lines change, the previous file is kept as.1,y/n,--dry-run.
A key pasted on the host it came from, or naming another host thanpeer.hostname, is
refused. No more editing the thumbprints by hand. -
ripcord statussays the listener is running, in oneLISTENERline with the build
its process runs, instead of leaving a missing block to mean it. After an update without a
restart, a second line namesripcord service restart. -
ripcord service startandripcord service stop.startstarts a stopped listener
without confirmation and leaves a running one alone.stopasksy/n— the
other host cannot read this one until the listener starts again — and waits for Windows to
report it stopped. -
ripcord serviceshows the build the running listener runs, from a record the listener
writes when it starts (logs\listener-process.txt), believed only when its process id is
the one Windows gives. Afterripcord updatewithout a restart it says the process still
runs the old build and ends withripcord service restart. A listener started before this
version has no record:unknownuntil it restarts.
Changed
- The samples' placeholder thumbprints are refused by name. A
listenerblock still
carryingAAAA1111…or1111AAAA…used to load, and the pair channel then failed as a
SILENTpeer, which reads as a network fault. Aripcord.yamlwhose listener is enabled
with them no longer loads: put the thumbprints of the two hosts' certificates, or set
enabled: false, behind which they are ignored.
Fixed
- A refusal no longer blames Hyper-V when Hyper-V was not involved. Every failure line
began "cannot read the local Hyper-V state", includingcheck-updaterefusing because
checking was off. Only a failed Hyper-V read says so now. updates.checkandupdates.installrefusals say how to switch them on:
set updates.check: true in ripcord.yaml, and likewise for install.- Failure and refusal lines wrap at 75 columns, so a reason from Windows or the network
no longer runs off a 1024×768 console.
Before you update
Update both hosts, one after the other, and finish. Ripcord's failover sequences are encoded
in the binary and they span two hosts, so a pair running two versions would execute half a
sequence written by each. A planned failover and a failback therefore refuse on a version
mismatch rather than warning — which means a pair left half-updated cannot be moved electively.
The window between the two hosts is the one time the tool will not work.
The one exception is failover --scenario unplanned, which runs entirely on the host it is
typed on. There is no half for the other binary to execute, and refusing a disaster failover for
want of a version string the dead host never published would fail at the one thing this tool
exists for.
Do it when nothing is on fire, not during an incident.
Verify what you are installing
There is no Authenticode signature yet; it needs a paid certificate. Until then the SHA-256
checksum published beside the binary is the only way to tell that the file on the host is the
file that was built. Check it before replacing anything.
(Get-FileHash ripcord.exe -Algorithm SHA256).Hash.ToLower()Updating
ripcord update installs a newer release on the host it is run on. It is off unless
updates.install says so, it asks y/n with Enter declining, and it refuses any release
whose detached signature does not verify against the key compiled into the running binary.
Copying the .exe by hand still works and is still supported; fully offline operation stays
possible.
This reverses what earlier versions of this file said. The objection was sound and has not gone
away: a binary that replaces itself, with Hyper-V privileges, on both hosts of a disaster
recovery pair, from the internet, is a supply chain reaching past the controls the rest of the
tool is built around. What changed is that the path now has a trust root that is not the release
page — the signature is checked against a pinned key, and an unverifiable release is refused
before anything is touched. What has not changed is that the signing key lives in the
release workflow, so an account with write access to the repository can still sign; see
SECURITY.md.
Update one host at a time, and update both. While the two differ, a failover spanning them
is refused — the sequences are encoded in the binary and one executed half by each version is
the error nobody recovers from at 3 a.m. The command says which host to run next.
The full procedure is in docs/RELEASING.md.