Repository navigation
v0.3.0
clusterctl v0.3.0
This release makes clusterctl more careful. Exit codes now mean what the manual says they mean, text from nodes, BMCs and infrastructure hosts can no longer take over your terminal, the MCP server holds an agent to the site and the plan it showed, and the Redfish client and the tunnel stop trusting what they should not. No command or flag is new. Some of these fixes change exit codes or refuse configuration that v0.2.0 accepted, which a minor release before 1.0.0 may do; Upgrading from v0.2.0 lists them.
Exit codes and interrupts
- Usage errors exit 2. An unknown flag, a bad flag value, a wrong number of arguments or an unknown command used to exit 1, which reads as "some nodes are unhealthy". A misspelt subcommand, such as
clusterctl slurm node drian, printed help and exited 0; it is now refused, with a suggestion.helpandcompletionfollow the same rule. - An interrupt exits 130, whichever command it stopped. The first Ctrl-C now also ends a confirmation or password prompt, and a second one ends the process. ssh is sent SIGTERM rather than killed, so it can restore the terminal, and no further node is started once a fan-out is cancelled.
- A remote exit 255 is no longer read as an unreachable host. A remote command that exits 255 is reported as 254, so 255 comes only from ssh itself.
- A command that fails on an infrastructure host exits 1, not 3, and the error quotes the first line the host wrote.
- A panic in one fan-out target is that target's failure; the others finish and are reported.
Terminal safety
Everything a node, a BMC, Slurm, a PXE host, a DHCP file or a group source prints now goes through one escaper, and so do error messages. Bidirectional controls, line and paragraph separators and bytes that are not UTF-8 are shown as escapes, so output cannot be reordered, forged as another node's line, or write to your clipboard. -o json is unchanged.
MCP server
- A plan is bound to the cluster, hosts and commands it showed.
apply_planrefuses it if the configuration has since changed, and the question is built from what applying would do now, with the server's own summary last, next to the answer. - Every refusal is written to the audit log. A plan that cannot be recorded is not offered, and an apply that cannot be recorded sends nothing.
read_commandreaches only the site's hosts, ignoresCLUSTERCTL_NODES, pins--fanout, may only shorten--timeout, stops at its output limit, and no longer offershelporcompletion.- ssh and password helpers never prompt without a terminal (
BatchMode=yes), so a call cannot hang on a password prompt. - A panicking tool becomes a failed call instead of ending the server.
BMCs and Redfish
hostkeyanddns lookup --bmc,describe_nodesandbmc webnow reach thebmcAddressthe inventory records, like every other BMC command.- Resets are checked against the vendor profile's
resetTypes, and against@Redfish.ActionInfowhere the firmware lists its reset types only there. - Redirects are refused rather than followed, so credentials cannot be sent to another scheme or port.
- A first-use certificate pin is recorded only once, even when several connections race.
Other fixes
tunnelruns sshuttle with the site's generated ssh configuration, host keys and account, stops only the process that is the profile's tunnel, and refuses an exclude that expands to nothing.--forcenames every protected or unknown host it lets through, also with-y.boot set --dry-run,provision reinstall --dry-runanddoctor --remote --dry-runrun the same read-only checks as the real run.node groups NODEandbmc web NODEselect their argument like every other node argument.- The output kept of a remote command is bounded at 32 MiB per stream; a result that was cut off is a failure.
- The group list of an exec source is cached like its memberships.
Upgrading from v0.2.0
The command line gains nothing and loses nothing, and the schema is still clusterctl/v1alpha1. Check these before you upgrade:
- Scripts that read exit codes. A usage error is now 2 rather than 1 (or 0, for a misspelt subcommand); an interrupt is 130; a remote exit of 255 arrives as 254; a command failing on an infrastructure host is 1 rather than 3.
- Node names with capitals in a
NodeInventoryare refused on load. Write them in lower case; host names are not case sensitive, and such nodes were never found before. fanout.maxof 0 or less, in any layer,CLUSTERCTL_FANOUT,--setor--fanout, is refused. It used to become 16.- A
Clusterwithoutinventoriesnow uses only its own site's inventories. With several sites loaded it is refused and told to list them. - A tunnel profile that sets
--ssh-cmd,--remote,--pidfileor--daemonin its options is refused, and one whose exclude expands to nothing is refused too.
clusterctl config validate reports the configuration changes, with file and line.
mise offers a release only once it is 24 hours old. To take this one sooner, name it:
$ mise use -g github:GSI-HPC/clusterctl@0.3.0Download the binary for your platform below, or:
$ go install github.com/GSI-HPC/clusterctl/cmd/clusterctl@v0.3.0Changelog
Features
- 8a4a8c6: feat(app): keep a runner for read-only lookups in a dry run (@claude)
- 08a09ee: feat(cli): add boot grub unset, and say that a GRUB link persists (@claude)
- c77b7e7: feat(hostname): decide whether a name may be used as a host (@claude)
Fixes
- cfd1c07: fix(apis): drop the fields that nothing reads (@claude)
- 37a396d: fix(app): apply --fanout and command line paths the way config explain reports them (@claude)
- 8f00f0b: fix(app): never read another site's inventory for a cluster that names none (@claude)
- 863a542: fix(app): read remote files for real in a dry run, cached per target (@claude)
- c2fba77: fix(app): refuse node names that are not host names when selected (@claude)
- 90a7af2: fix(app): report a failed command on a role as a failure, with what it said (@claude)
- c17f235: fix(app): select machines, not spellings (@claude)
- 129d3a4: fix(bmc): gate bmc forget, and forget only the pins that are there (@claude)
- 2d6548d: fix(bmc): reach the service processor the inventory records, by any spelling (@claude)
- ced27a8: fix(bmc): read the ping sweep carefully and report a failed sweep as such (@claude)
- d6cbf58: fix(bmc): report every node honestly, grouped by account, and batch cycles (@claude)
- 460722b: fix(cli): ask Slurm about every node when the host list is too long (@claude)
- bc62b43: fix(cli): ask Slurm in a dry run as the real run does (@claude)
- a7f7118: fix(cli): ask before boot sync pulls the boot configurations (@claude)
- d882141: fix(cli): ask the name server itself in the dns commands (@claude)
- c2f3acb: fix(cli): bound boot log --lines (@claude)
- 2c7013b: fix(cli): check the PXE host and the roles in a dry run too (@claude)
- cc7bbf9: fix(cli): complete node sets without running group commands (@claude)
- f04bec8: fix(cli): complete with --config and -n on the command line (@claude)
- 82084d4: fix(cli): copy through the fan-out, into a directory per node (@claude)
- f54a3c1: fix(cli): count every busy Slurm state, and refuse the unknown ones (@claude)
- 4654fce: fix(cli): describe a node under the name the inventory gives it (@claude)
- accbeac: fix(cli): escape node output through output.EscapeText (@claude)
- 6a59f37: fix(cli): escape the error report() prints (@claude)
- 74f23aa: fix(cli): exit 2 for usage errors in cobra's help and completion commands (@claude)
- fd7a99d: fix(cli): exit 2 for usage errors, and refuse an unknown subcommand (@claude)
- adba855: fix(cli): exit 3 or 130 from a fan-out, and show each node's own error (@claude)
- 9d77bb2: fix(cli): give node hw an entry for every failed node (@claude)
- 66f1c16: fix(cli): give the Slurm job check its own override, --lose-jobs (@claude)
- ee87bdc: fix(cli): keep config init from shadowing a configuration on the search path (@claude)
- 7c4aed6: fix(cli): list every node in provision status, and fail for what it cannot read (@claude)
- 1519cc4: fix(cli): make doctor check what it says it checks (@claude)
- 1143349: fix(cli): quote the dhcp log path and bound dhcp log and capture (@claude)
- 14b68a8: fix(cli): refuse a node set given both as an argument and with -n (@claude)
- e4ffcbb: fix(cli): refuse a second name in login (@claude)
- 24bbe51: fix(cli): refuse an empty or repeated -n instead of guessing (@claude)
- ffbda36: fix(cli): refuse what Slurm did not report, and check Redfish resets (@claude)
- 1726140: fix(cli): report fabric ports, uplinks and GUIDs as they are (@claude)
- d9cf01d: fix(cli): report the nodes cinc show could not read (@claude)
- c077cb7: fix(cli): resolve a reinstall before it asks, check Slurm, and disarm what a failure armed (@claude)
- d82a1fe: fix(cli): resolve and check boot links before asking, and report each node (@claude)
- 2dbe10a: fix(cli): resolve the BMC the inventory names in dns lookup --bmc (@claude)
- 13fd586: fix(cli): resolve the node that node groups NODE names (@claude)
- 89141af: fix(cli): run the exec gate on every run, and read the command only after -- (@claude)
- 774a89f: fix(cli): say when the Slurm job check is skipped (@claude)
- ce1969a: fix(cli): scan and remove the host keys of the BMC the inventory names (@claude)
- 4598692: fix(cli): secrets push decrypts first, names unreachable nodes (@claude)
- 8b889bf: fix(cli): select the node bmc web opens the service processor of (@claude)
- 20444b9: fix(cli): send Redfish requests through one fan-out, and recover a panic in it (@claude)
- 3b595e3: fix(cli): show a failed node's own error in exec, with control characters escaped (@claude)
- 8f9b2d1: fix(cli): split hca config into get and set, and set every adapter (@claude)
- 15e7232: fix(cli): stop a jq program of version and config init with the command (@claude)
- 22e0ce5: fix(cli): write the cinc file quoted and read it without a shell (@claude)
- 1261482: fix(config): check overrides against the schema and merge them key by key (@claude)
- 46402d9: fix(config): drop the fan-out offload settings nothing reads (@claude)
- dc47ab4: fix(config): pin the scaffold's cluster to its site's inventory and quote every name (@claude)
- e5134b9: fix(config): refuse BMC settings the code does not keep (@claude)
- a5eff54: fix(config): refuse a fanout.max below one instead of reading it as 16 (@claude)
- ee7eef1: fix(config): refuse configuration someone else could have written (@claude)
- a387b2e: fix(config): refuse role hosts and users that ssh would read as something else (@claude)
- a24aa4b: fix(config): refuse to run without a private state and cache directory (@claude)
- 5ed1df2: fix(config): report a --config or CLUSTERCTL_CONFIG entry that is missing (@claude)
- 75eebeb: fix(config): turn the Slurm job check on by default (@claude)
- 91f5259: fix(credentials): read a password once however many lookups run at once (@claude)
- b491d8b: fix(credentials): run a password helper named by a path from the site (@claude)
- e719895: fix(dhcp): take the boot address only from the node's own declaration (@claude)
- b86a996: fix(dhcp): tokenise dhcpd.conf and follow include (@claude)
- a0bbcdb: fix(fanout): group --dedup answers by how the nodes ended (@claude)
- 851a5ae: fix(fanout): report a target GroupByOutput cannot group (@claude)
- fe27550: fix(fanout): turn a panic in a fan-out worker into that target's failure (@claude)
- 5472bf8: fix(fileutil): share locks with the group, keep mode, group and links (@claude)
- 6723a26: fix(groups): cache the group list of an exec source (@claude)
- f819c37: fix(groups): escape a failing source's message through output.EscapeCell (@claude)
- 7e3f48b: fix(groups): scope the group cache and fail closed when a source fails (@claude)
- 58594ed: fix(hostkeys): keep comments and markers when rewriting the host key file (@claude)
- 4e90f9f: fix(hostkeys): report the jump host's reason when the write fails first (@claude)
- 01a797e: fix(hostkeys): scan ssh-rsa hosts, through their jump host, in parallel (@claude)
- 4407f4b: fix(inventory): keep an entry's rack attribute, and read racks from it (@claude)
- 9728e2f: fix(inventory): keep the parse error of a malformed identifier (@claude)
- 7a8c6f5: fix(inventory): refuse a node name written with capitals (@claude)
- d711aa0: fix(inventory): refuse shared or malformed machine identifiers (@claude)
- 640ecff: fix(inventory): refuse two spellings of one host (@claude)
- 50437cd: fix(ipmi): report every processor, and only known answers as success (@claude)
- 5c85776: fix(mcp): bind a plan to its cluster and record every outcome (@claude)
- fa3d234: fix(mcp): hold read_command to the site's hosts and no default node set (@claude)
- e5ecde2: fix(mcp): keep cobra's help and completion out of read_command (@claude)
- 4a7ec36: fix(mcp): pin --fanout and cap --timeout in read_command (@claude)
- e4f3df0: fix(mcp): put the summary last in the confirmation question (@claude)
- 2a3f2b6: fix(mcp): recover a panicking tool and stop read_command at its bound (@claude)
- f7ca04e: fix(mcp): show the BMC the commands reach in describe_nodes (@claude)
- a5397a6: fix(naming): never name a node's service processor after the node (@claude)
- edbf63d: fix(nodeset): fold Hostlist along one dimension and without steps (@claude)
- 0e7835e: fix(nodeset): keep the spelling every host was given (@claude)
- 50c5088: fix(nodeset): refuse an oversized expression before expanding it (@claude)
- abfee91: fix(nodeset): reject a set operator without an operand (@claude)
- f2748c6: fix(output): compile queries when -o is read and bound jq by the context (@claude)
- 71a1045: fix(output): escape control characters in tables and names (@claude)
- 2ee477b: fix(output): keep integers in -o yaml and quote what a reader would misread (@claude)
- 30ef6a5: fix(redfish): check resets against the vendor profile and ActionInfo (@claude)
- 9f5dc1a: fix(redfish): exit 3 when a processor cannot be reached or turns the account away (@claude)
- b2a8ca5: fix(redfish): record a first-use pin only once (@claude)
- 4ded01f: fix(redfish): refuse a client that can neither verify nor pin (@claude)
- 1c8978d: fix(redfish): refuse redirects instead of following them (@claude)
- ae466c3: fix(redfish): send a request only to a host that carries nothing else (@claude)
- f66907c: fix(release): verify the signer against a list the tag cannot change (@claude)
- 5d9fe46: fix(safety): decide protected hosts by machine, and refuse nodes nobody knows (@claude)
- e2064f2: fix(safety): make confirmAbove 0 ask for the count every time (@claude)
- 34b6249: fix(safety): name the hosts --force lets through (@claude)
- 36c9b07: fix(secrets): keep decrypted values out of errors (@claude)
- 326b351: fix(secrets): trust only the configured key types, and age first (@claude)
- 73d749a: fix(slurm): check node sets and reasons before a change, read output safely (@claude)
- 002d069: fix(transport): bound the output kept of a remote command (@claude)
- b2bfeec: fix(transport): name the fields of a target and carry the error in results (@claude)
- 82fdd76: fix(transport): put the ssh destination after -- (@claude)
- 5888974: fix(transport): stop ssh gently, and start no target once cancelled (@claude)
- 0c10d1b: fix(transport): tell a remote exit 255 from ssh's own (@claude)
- 687a9e1: fix(transport): trust one host key file, one generated file per configuration (@claude)
- b12c9c5: fix(tunnel): connect sshuttle with the generated ssh configuration (@claude)
- fb3f534: fix(tunnel): refuse an exclude that expands to nothing (@claude)
- 8e628a3: fix(tunnel): stop only the process that is the profile's tunnel (@claude)
- c40aeae: fix(version): report devel and the revision for a build from a checkout (@claude)
- 6459cc3: fix: escape what infrastructure hosts print, through the one helper (@claude)
- 273853f: fix: exit 130 when an interrupt stops a command (@claude)
- dbff5f8: fix: let a second Ctrl-C end the process, and the first end a prompt (@claude)
- 56b66a7: fix: never let ssh or a password helper prompt without a terminal (@claude)
Documentation
- 4062c1e: docs(config): say how overrides, duplicates and config init keep protections (@claude)
- 878a014: docs(migration): map cluster-reboot-node to an action that reboots (@claude)
- 11349a7: docs(output): read the exit code before anything else consumes it (@claude)
- 63d0f76: docs(readme): name the macOS configuration directory (@claude)
- a8f523f: docs(safety): state the dry-run rule for read-only lookups (@claude)
- 7067836: docs(slurm): describe the checks before drain, resume and accounting changes (@claude)
- 1d2970e: docs: describe escaping, the JSONPath grammar and the result fields (@claude)
- 139a1cd: docs: describe how a reinstall resolves, checks Slurm and disarms after a failure (@claude)
- 95730b8: docs: describe how exec reads its arguments, checks and reports (@claude)
- 839573f: docs: describe the Slurm job check as it now works (@claude)
- c0e486f: docs: describe the fabric, doctor, dns and completion changes (@claude)
- 53180b6: docs: describe the generated ssh configuration as it now is (@claude)
- 2b5af90: docs: describe the persistent boot link and the checks before a boot link (@claude)
- b8a65df: docs: say how inventory entries name hosts, racks and addresses (@claude)
- ce541a9: docs: say that errors are escaped, and which Unicode controls are (@claude)
- 979f89d: docs: say that hostkey and dns --bmc reach the inventory's bmcAddress (@claude)
- 4b8698c: docs: say what the MCP server checks at apply time and what read_command reaches (@claude)
- d870845: docs: say which sops keys are trusted and what secrets push does (@claude)
- 3db382c: docs: stop offering @idle and @drained, which no group source provides (@claude)
Other
- 85481fe: perf(inventory): build the set of node names once (@claude)
- 428f475: refactor(transport): leave WaitDelay of a copy to the command builder (@claude)
- 2c9201f: refactor(transport): remove Client.CheckConfig, which nothing called (@claude)
The manual covers every command.
Verify a download against checksums.txt.