CiruStrixLink 0.3.0
CiruStrixLink 0.3.0
Version 0.3.0 turns the embedded console into a practical two-machine operator
surface while preserving CiruStrixLink's fail-closed transport behavior.
Highlights
- Redesigned Connection Setup as a guided two-host workflow.
- Replaced Launch Settings with a real Launch page and moved it to the third
tab. - Added coordinated configure, load, unload, readiness, and rollback handling
for GLM5.3 Flash CIRU STRIX IU4. - Added live rank, process, context, KV allocation, host-wide unified RAM, and
available-memory telemetry for both machines. - Fixed Fast-mode reporting so Overview, Diagnostics, and Launch use the same
scoped privileged truth when model control is enabled. - Added copyable permission setup for packaged NixOS and generic Linux.
- Added a built-in, root-only generic Linux helper with exact sudoers commands.
Safe paired lifecycle
Model control is opt-in and requires a shared peer token, complementary fixed
ranks, fixed USB4 IPv4 peers, a model frontend URL, a loopback-only console,
and narrow root-owned helpers on both hosts.
A load verifies both ranks and the NHI lease, writes the same selected profile
to both stopped ranks, starts rank 0 followed by rank 1, and waits until the
frontend serves the exact context. Failure stops both ranks again and reports
any incomplete rollback. Unload stops rank 1 before rank 0.
CiruStrixLink refuses to load this deployment while qwen-main.service or the
portable GLM service is active. It never enables, disables, masks, or replaces
those services, and it never changes a context while the model is running.
Desktop access
Run the console on one model host and the USB4-bound agent on the other. From a
desktop, forward the console's loopback port:
ssh -N -L 7749:127.0.0.1:7749 USER@CONSOLE_MODEL_HOSTOpen http://127.0.0.1:7749/#/launch. The desktop needs no StrixLink binary or
model-host privileges.
See the browser-console instructions,
GLM deployment guide,
and security model.
256K profile
The released 256K profile configures 262,272 tokens with an 8 GiB KV allocation
per rank and 8 GiB of host prefix staging per rank. It is a single-request,
experimental high-context profile. The UI reports host-wide unified memory so
operators can see the actual remaining headroom rather than treating the KV
allocation as total model memory.
Validation
- Complete Go tests passed.
go vetpassed.- Embedded JavaScript syntax and whitespace checks passed.
- Static Linux amd64 release build passed.
- Live read-only validation found both 256K ranks loaded, the frontend serving
the exact 262,272-token context, and Fast modein_usewith qualified HopID
9/9. The running model was not stopped, restarted, or reconfigured.
See CHANGELOG.md
for the complete change list.