Repository navigation
Releases: GSI-HPC/clusterctl
Release list
v0.4.0
clusterctl v0.4.0
This release shows how far a long command has got, runs on many hosts at once what still went one host at a time, and decrypts Secret documents with the sops command, which takes the binary to a third of its size. It also removes the configuration keys nothing read and the -o jsonpath format, and gives the exit code of a command on many hosts one rule. Some of these changes refuse configuration, command lines or Secret documents that v0.3.0 accepted, which a minor release before 1.0.0 may do; Upgrading from v0.3.0 lists them.
Progress
--progress, orCLUSTERCTL_PROGRESS, chooses what standard error shows while a command runs.auto, the default, draws a live tree of the steps and the hosts under way when standard error is a terminal and standard output does not go into a pipe, and nothing otherwise;ttyasks for the tree,counterfor a single line,plainfor lines a CI log can keep, andnonefor nothing. Standard output never carries progress, and without a terminal standard error holds what it held in v0.3.0, byte for byte, unlessplainis asked for.- A display that has run for a second or more ends with one summary line, such as
clusterctl: provision reinstall: failed in 18m03s: 478 ok, 2 failed. The summary and the plain lines are written for people and may change; programs read the event log. --progress-log FILE, orCLUSTERCTL_PROGRESS_LOG, appends every event to a file as a line of JSON, in format version 1, and creates the file with mode 0600. ATRACEPARENTin the environment gives the log its trace; neither it norTRACESTATEis handed on to the programs clusterctl runs.--progress ttyor--progress counterwhere neither can be drawn, and a--progress-logthat cannot be used, are usage errors (exit 2). The variables fail no command: they show nothing, or write no log, and say why in one line.- The MCP server sends progress notifications for its tool calls, and its audit log records each call's trace.
The Progress page of the manual describes the displays and the event log.
Many hosts at once
dns lookup,provision status,secrets push,doctor --remote,bmc powerthrough the ipmitool backend,fabric state, and the MCP toolsdescribe_nodesandplan_changenow work on many hosts at once rather than one after another.- Two new settings bound them:
services.dns.maxConcurrent, 16 by default, andbmc.ipmi.maxConcurrent, 8.--fanouton the command line lowers these and the Redfish limit but never raises them, and is handed to ipmipower;CLUSTERCTL_FANOUTandfanout.maxleave them alone. doctor --remoteasks each role in one session. A role whose account logs in but runs nothing, such as one with a nologin shell, now fails.- The MCP server runs at most two tool calls at once; the others wait, reported as
waiting for another tool call.
Secret documents are decrypted by the sops command
A workstation that reads a sops-encrypted Secret needs sops 3.10.0 or later, in PATH or named by workstation.sopsBinary or CLUSTERCTL_SOPS_BINARY; clusterctl doctor checks for it. Commands that use no secret do not need sops. The linux/amd64 binary shrinks from 50.5 MB to 17.3 MB, and fixes to sops and its cloud SDKs now come with your sops update rather than with a clusterctl release. clusterctl still makes every check it made before sops runs, and hands sops only the bytes it has checked. It now also refuses sops metadata with a field it does not know, and master keys both inside and outside key_groups, where sops would ignore the groups.
Exit codes
A command on many hosts now exits with the worst of its hosts' codes, in one order: 130, then 3, then 2, then 1.
secrets pushwith one node refused and one unreachable exits 3 (was 1).exec,copy,cinc,fabric hcaandnode hardwareexit 2 (was 1) when a node's configuration is at fault.cincwith one node interrupted and one unreachable exits 130 (was 3).hostkey verify,hostkey refresh,bmc ping,dns lookupanddns aliaseskeep rules of their own, which the exit code reference now lists.secrets pushrefuses two secrets for one target (exit 2), and exits 3 when a node that failed could not be reached for any of its secrets.- A remote command is also stopped locally, with exit 3, once its timeout, the grace period and the time it may take to reach the host have passed, so a host that hangs without dropping the connection no longer holds a command up.
- A command on an infrastructure host whose output was cut off at the limit exits 3 (was 1), as Slurm's commands already did.
- An
ageFilecredential whose file or identities cannot be read exits 2 (was 1), as anageFilesecret does.
Configuration
- Keys nothing read are gone, and refused on load as unknown:
bmc.ipmi.passwordTransport,bmc.pdu.credential,fanout.connectTimeout,services.cinc.archivePath,services.cinc.baseCookbook,services.cinc.rolesPath,services.fabric.guidFormat,services.http.root,services.http.baseUrl,services.mail.*andworkstation.pager.CLUSTERCTL_PAGERgoes with the last. - An explicit
""or0for a key the defaults set, such asservices.dhcp.configPathorbmc.redfish.maxConcurrent, is refused; it used to be read as the default.services.tftp.rootandservices.tftp.logPathmay no longer be empty. - The context, tunnel and PDU users must be portable POSIX user names.
alice@EXAMPLE.ORGandsvc$now failconfig validate, with their line, where they used to validate and then fail every connection, andconfig init --userrefuses them too..aliceis now accepted, as connecting already did. - A ProxyJump host that is not a role is checked as a host name, so
gw_1.example.orgis refused.
Boot
- New:
boot grub logshows what the TFTP server,in.tftpd,tftpd,atftpdordnsmasq-tftp, wrote intoservices.tftp.logPath. boot grub setrefuses a target that is not a file underservices.tftp.root(exit 2, with--dry-runtoo), and writes the link relative, so that a TFTP server confined to its root, such asin.tftpd -s, finds it.
Output
-o jsonpath=…is gone: it exits 2, as an unknown format.-o jq=…does what its templates did, and the output guide gives the jq for each.boot log,boot sync,dhcp logandfabric countershonour-o jsonand-o yaml, as a list of lines.- Table columns that hold CJK characters or emoji line up.
- A failed remote command reads
<program> on <host> exited N: <what it said>everywhere. slurm job summary -o jsonrows have the key order of the MCP tool's, and credential names and ties in the job summary come in a fixed order.fabric stateprints the ports that answered when the fabric stops early, and asks the fabric in a dry run too.
Other fixes
- A Ctrl-C at a password prompt ends the reading of credentials:
provision statusused to ask again for every node, and could leave the terminal without echo. A credential lookup that failed is remembered rather than tried again for every node. - A worker's panic is written after a password prompt another node asks, never into it.
- Host key scans stop once interrupted.
cinc rungets its 30 minutes, orfanout.commandTimeoutwhen that is longer.- Idle Redfish connections time out, and are closed after a fan-out.
- scp's progress meter is shown only for a single transfer.
- ipmitool lines longer than 400 bytes are cut and end in
…. tunnel start,tunnel stopandconfig use-contextcomplete their names when--configis given.mcp serve --setis parsed as every other command's is.- A Redfish client or an IPMI backend printed in a debug line no longer shows its password.
describe_nodesno longer stops reading groups at the first node that fails.- A path such as
~bob/xis no longer read as$HOMEfollowed bybob/x. - A cached remote file with a modification time in the future is no longer trusted for ever. The cache format changed, so old entries read as misses once.
For Go programs that import nodeset
nodeset.Resolver now holds only what parsing calls, Resolve and All. List and DefaultSource moved to the new, optional Lister interface, which MapResolver implements. A resolver written for v0.3.0 still satisfies Resolver; code that called List or DefaultSource through a Resolver asserts Lister now. Folded output and error messages are unchanged. The package is to move to a module of its own, github.com/GSI-HPC/go-nodeset; a later release will say when.
Upgrading from v0.3.0
The command line gains --progress, --progress-log and boot grub log, and loses -o jsonpath; the schema is still clusterctl/v1alpha1. Check these before you upgrade:
- sops on every workstation that reads a Secret document, 3.10.0 or later, in
PATHor named byworkstation.sopsBinary. - Every workstation before the site's configuration. v0.3.0 refuses a key it does not know, so upgrade everyone before a site sets
services.dns.maxConcurrentorbmc.ipmi.maxConcurrent. - The removed keys, in every layer, and explicit empty values for keys the defaults set.
- User names and jump hosts that the stricter rules refuse.
- Scripts that read exit codes or use
-o jsonpath. See Exit codes and Output. xargson the IPMI gateway and the fabric host. The ipmitool backend andfabric statenow runxargs -0 -Pthere, which a BusyBox built without those options lacks.- GRUB targets outside
services.tftp.root, whichboot grub setrefuses. - Secret documents whose sops metadata has a field clusterctl does not know, or master keys both inside and outside
key_groups.
`clusterctl config validat...
v0.3.0
clusterctl v0.3.0
This release makes clusterctl more careful. Exit codes now mean what the manual says they mean, text from nodes, BMCs and infrastructure hosts can no longer take over your terminal, the MCP server holds an agent to the site and the plan it showed, and the Redfish client and the tunnel stop trusting what they should not. No command or flag is new. Some of these fixes change exit codes or refuse configuration that v0.2.0 accepted, which a minor release before 1.0.0 may do; Upgrading from v0.2.0 lists them.
Exit codes and interrupts
- Usage errors exit 2. An unknown flag, a bad flag value, a wrong number of arguments or an unknown command used to exit 1, which reads as "some nodes are unhealthy". A misspelt subcommand, such as
clusterctl slurm node drian, printed help and exited 0; it is now refused, with a suggestion.helpandcompletionfollow the same rule. - An interrupt exits 130, whichever command it stopped. The first Ctrl-C now also ends a confirmation or password prompt, and a second one ends the process. ssh is sent SIGTERM rather than killed, so it can restore the terminal, and no further node is started once a fan-out is cancelled.
- A remote exit 255 is no longer read as an unreachable host. A remote command that exits 255 is reported as 254, so 255 comes only from ssh itself.
- A command that fails on an infrastructure host exits 1, not 3, and the error quotes the first line the host wrote.
- A panic in one fan-out target is that target's failure; the others finish and are reported.
Terminal safety
Everything a node, a BMC, Slurm, a PXE host, a DHCP file or a group source prints now goes through one escaper, and so do error messages. Bidirectional controls, line and paragraph separators and bytes that are not UTF-8 are shown as escapes, so output cannot be reordered, forged as another node's line, or write to your clipboard. -o json is unchanged.
MCP server
- A plan is bound to the cluster, hosts and commands it showed.
apply_planrefuses it if the configuration has since changed, and the question is built from what applying would do now, with the server's own summary last, next to the answer. - Every refusal is written to the audit log. A plan that cannot be recorded is not offered, and an apply that cannot be recorded sends nothing.
read_commandreaches only the site's hosts, ignoresCLUSTERCTL_NODES, pins--fanout, may only shorten--timeout, stops at its output limit, and no longer offershelporcompletion.- ssh and password helpers never prompt without a terminal (
BatchMode=yes), so a call cannot hang on a password prompt. - A panicking tool becomes a failed call instead of ending the server.
BMCs and Redfish
hostkeyanddns lookup --bmc,describe_nodesandbmc webnow reach thebmcAddressthe inventory records, like every other BMC command.- Resets are checked against the vendor profile's
resetTypes, and against@Redfish.ActionInfowhere the firmware lists its reset types only there. - Redirects are refused rather than followed, so credentials cannot be sent to another scheme or port.
- A first-use certificate pin is recorded only once, even when several connections race.
Other fixes
tunnelruns sshuttle with the site's generated ssh configuration, host keys and account, stops only the process that is the profile's tunnel, and refuses an exclude that expands to nothing.--forcenames every protected or unknown host it lets through, also with-y.boot set --dry-run,provision reinstall --dry-runanddoctor --remote --dry-runrun the same read-only checks as the real run.node groups NODEandbmc web NODEselect their argument like every other node argument.- The output kept of a remote command is bounded at 32 MiB per stream; a result that was cut off is a failure.
- The group list of an exec source is cached like its memberships.
Upgrading from v0.2.0
The command line gains nothing and loses nothing, and the schema is still clusterctl/v1alpha1. Check these before you upgrade:
- Scripts that read exit codes. A usage error is now 2 rather than 1 (or 0, for a misspelt subcommand); an interrupt is 130; a remote exit of 255 arrives as 254; a command failing on an infrastructure host is 1 rather than 3.
- Node names with capitals in a
NodeInventoryare refused on load. Write them in lower case; host names are not case sensitive, and such nodes were never found before. fanout.maxof 0 or less, in any layer,CLUSTERCTL_FANOUT,--setor--fanout, is refused. It used to become 16.- A
Clusterwithoutinventoriesnow uses only its own site's inventories. With several sites loaded it is refused and told to list them. - A tunnel profile that sets
--ssh-cmd,--remote,--pidfileor--daemonin its options is refused, and one whose exclude expands to nothing is refused too.
clusterctl config validate reports the configuration changes, with file and line.
mise offers a release only once it is 24 hours old. To take this one sooner, name it:
$ mise use -g github:GSI-HPC/clusterctl@0.3.0Download the binary for your platform below, or:
$ go install github.com/GSI-HPC/clusterctl/cmd/clusterctl@v0.3.0Changelog
Features
- 8a4a8c6: feat(app): keep a runner for read-only lookups in a dry run (@claude)
- 08a09ee: feat(cli): add boot grub unset, and say that a GRUB link persists (@claude)
- c77b7e7: feat(hostname): decide whether a name may be used as a host (@claude)
Fixes
- cfd1c07: fix(apis): drop the fields that nothing reads (@claude)
- 37a396d: fix(app): apply --fanout and command line paths the way config explain reports them (@claude)
- 8f00f0b: fix(app): never read another site's inventory for a cluster that names none (@claude)
- 863a542: fix(app): read remote files for real in a dry run, cached per target (@claude)
- c2fba77: fix(app): refuse node names that are not host names when selected (@claude)
- 90a7af2: fix(app): report a failed command on a role as a failure, with what it said (@claude)
- c17f235: fix(app): select machines, not spellings (@claude)
- 129d3a4: fix(bmc): gate bmc forget, and forget only the pins that are there (@claude)
- 2d6548d: fix(bmc): reach the service processor the inventory records, by any spelling (@claude)
- ced27a8: fix(bmc): read the ping sweep carefully and report a failed sweep as such (@claude)
- d6cbf58: fix(bmc): report every node honestly, grouped by account, and batch cycles (@claude)
- 460722b: fix(cli): ask Slurm about every node when the host list is too long (@claude)
- bc62b43: fix(cli): ask Slurm in a dry run as the real run does (@claude)
- a7f7118: fix(cli): ask before boot sync pulls the boot configurations (@claude)
- d882141: fix(cli): ask the name server itself in the dns commands (@claude)
- c2f3acb: fix(cli): bound boot log --lines (@claude)
- 2c7013b: fix(cli): check the PXE host and the roles in a dry run too (@claude)
- cc7bbf9: fix(cli): complete node sets without running group commands (@claude)
- f04bec8: fix(cli): complete with --config and -n on the command line (@claude)
- 82084d4: fix(cli): copy through the fan-out, into a directory per node (@claude)
- f54a3c1: fix(cli): count every busy Slurm state, and refuse the unknown ones (@claude)
- 4654fce: fix(cli): describe a node under the name the inventory gives it (@claude)
- accbeac: fix(cli): escape node output through output.EscapeText (@claude)
- 6a59f37: fix(cli): escape the error report() prints (@claude)
- 74f23aa: fix(cli): exit 2 for usage errors in cobra's help and completion commands (@claude)
- fd7a99d: fix(cli): exit 2 for usage errors, and refuse an unknown subcommand (@claude)
- adba855: fix(cli): exit 3 or 130 from a fan-out, and show each node's own error (@claude)
- 9d77bb2: fix(cli): give node hw an entry for every failed node (@claude)
- 66f1c16: fix(cli): give the Slurm job check its own override, --lose-jobs (@claude)
- ee87bdc: fix(cli): keep config init from shadowing a configuration on the search path (@claude)
- 7c4aed6: fix(cli): list every node in provision status, and fail for what it cannot read (@claude)
- 1519cc4: fix(cli): make doctor check what it says it checks (@claude)
- 1143349: fix(cli): quote the dhcp log path and bound dhcp log and capture (@claude)
- 14b68a8: fix(cli): refuse a node set given both as an argument and with -n (@claude)
- e4ffcbb: fix(cli): refuse a second name in login (@claude)
- 24bbe51: fix(cli): refuse an empty or repeated -n instead of guessing (@claude)
- ffbda36: fix(cli)...
v0.2.0
clusterctl v0.2.0
clusterctl now writes a first configuration for you, and the manual describes the release you install rather than what is on main.
What is new
clusterctl config init [DIR]writes the least configuration that resolves: aConfigwith one context, aSitewith a login node, aClusterand an emptyNodeInventory, one document to a file, with comments that say what to fill in.--site,--cluster,--domain,--loginand--userset the names. WithoutDIRthe files go where configuration is read from: the directory--configorCLUSTERCTL_CONFIGnames, or else your own configuration directory, usually~/.config/clusterctl. It writes only into an empty directory and never overwrites a file, and--dry-runlists the files without writing them. When no configuration is found, the error now points to it.clusterctl mcpdoes not offer it to an agent.- A manual for every release. The manual at the root now describes the latest release, not main. Each minor line keeps its own manual under
/vX.Y/, main's is under/dev/, and a switcher beside the title moves between them. The JSON Schemas that the$schemaline of a configuration file names now describe the latest release too.
Upgrading from v0.1.0
Nothing needs to change. No command, flag, exit code or configuration field was changed or removed, the schema is still clusterctl/v1alpha1, and a configuration written for v0.1.0 loads as before. config init is for starting a new configuration: it refuses a directory that already holds one.
mise offers a release only once it is 24 hours old. To take this one sooner, name it:
$ mise use -g github:GSI-HPC/clusterctl@0.2.0Download the binary for your platform below, or:
$ go install github.com/GSI-HPC/clusterctl/cmd/clusterctl@v0.2.0Changelog
Features
Documentation
- c41acea: docs(install): make the release download and mise steps work as written (@claude)
- 87891ba: docs(readme): lead with the tagline and a link to the manual (@claude)
- b39bd0d: docs(site): publish a manual for every release, the latest at the root (@claude)
- 2fc5ccb: docs(site): redesign the front page of the manual (@claude)
- fc9d181: docs: name GSI as the copyright holder (@claude)
Other
- 5bcceb4: build(deps): Bump google.golang.org/grpc in the go-modules group (@dependabot[bot])
The manual covers every command.
Verify a download against checksums.txt.
v0.1.0
clusterctl v0.1.0
The first release of clusterctl: one static binary in place of the 27 Bash scripts of the cluster toolkit. It reaches the hosts of a site, selects nodes with ClusterShell node set syntax, runs commands on them in parallel, drives their service processors, reinstalls them and administers Slurm, for as many clusters as the configuration describes.
What is in it
- Reach everything: one host key file kept in version control, a generated ssh configuration, and sshuttle profiles for the networks behind a gateway.
- Select and fan out: ClusterShell compatible node sets, groups from node attributes, tables or the workload manager, and bounded parallel execution that reports every node.
- Operate hardware: Redfish with a pinned certificate, FreeIPMI and ipmitool, rack power units and the InfiniBand fabric.
- Reinstall nodes: DHCP inspection, per-node PXE and GRUB boot paths, age or sops encrypted secrets streamed to the node, and the configuration management client.
- Administer Slurm: nodes and their drain reasons, the queue and the accounting database, accounts, users and fair share.
- Work with an agent:
clusterctl mcplets an AI agent read the cluster and plan changes, and nothing is sent until you confirm it. - Ask before changing anything: a command that changes something shows what it is about to do and asks first,
--dry-runsends nothing, and a protected host needs--force. - The node set engine as a library:
github.com/GSI-HPC/clusterctl/nodeset, for other Go programs.
Coming from the shell toolkit
There is no compatibility layer. doc/migration.md maps every command and setting to its counterpart, and examples/site/ in each archive is a complete configuration to start from.
Before you rely on it
- This is a 0.x release. The configuration schema is
clusterctl/v1alpha1, and a minor release may still change a command, a flag or a field; its notes will say so. - What needs real hardware or a real cluster (IPMI against a service processor, fabric diagnostics, a Slurm controller, a PXE boot) is tested up to the command handed to the transport, not against the hardware. Run it with
--dry-runon your site first.
Platforms and verification
Linux and macOS, on amd64 and arm64, statically linked. Every archive is listed in checksums.txt and carries a build provenance attestation:
$ gh attestation verify clusterctl_linux_amd64.tar.gz --repo GSI-HPC/clusterctlWith mise, mise use -g github:GSI-HPC/clusterctl installs it and checks both.
Download the binary for your platform below, or:
$ go install github.com/GSI-HPC/clusterctl/cmd/clusterctl@v0.1.0The manual covers every command.
Verify a download against checksums.txt.