Skip to content

CiruStrixLink 0.3.0

Choose a tag to compare

@ciru-ai ciru-ai released this 03 Sep 19:21
· 4 commits to master since this release

CiruStrixLink 0.3.0

Version 0.3.0 turns the embedded console into a practical two-machine operator
surface while preserving CiruStrixLink's fail-closed transport behavior.

Highlights

  • Redesigned Connection Setup as a guided two-host workflow.
  • Replaced Launch Settings with a real Launch page and moved it to the third
    tab.
  • Added coordinated configure, load, unload, readiness, and rollback handling
    for GLM5.3 Flash CIRU STRIX IU4.
  • Added live rank, process, context, KV allocation, host-wide unified RAM, and
    available-memory telemetry for both machines.
  • Fixed Fast-mode reporting so Overview, Diagnostics, and Launch use the same
    scoped privileged truth when model control is enabled.
  • Added copyable permission setup for packaged NixOS and generic Linux.
  • Added a built-in, root-only generic Linux helper with exact sudoers commands.

Safe paired lifecycle

Model control is opt-in and requires a shared peer token, complementary fixed
ranks, fixed USB4 IPv4 peers, a model frontend URL, a loopback-only console,
and narrow root-owned helpers on both hosts.

A load verifies both ranks and the NHI lease, writes the same selected profile
to both stopped ranks, starts rank 0 followed by rank 1, and waits until the
frontend serves the exact context. Failure stops both ranks again and reports
any incomplete rollback. Unload stops rank 1 before rank 0.

CiruStrixLink refuses to load this deployment while qwen-main.service or the
portable GLM service is active. It never enables, disables, masks, or replaces
those services, and it never changes a context while the model is running.

Desktop access

Run the console on one model host and the USB4-bound agent on the other. From a
desktop, forward the console's loopback port:

ssh -N -L 7749:127.0.0.1:7749 USER@CONSOLE_MODEL_HOST

Open http://127.0.0.1:7749/#/launch. The desktop needs no StrixLink binary or
model-host privileges.

See the browser-console instructions,
GLM deployment guide,
and security model.

256K profile

The released 256K profile configures 262,272 tokens with an 8 GiB KV allocation
per rank and 8 GiB of host prefix staging per rank. It is a single-request,
experimental high-context profile. The UI reports host-wide unified memory so
operators can see the actual remaining headroom rather than treating the KV
allocation as total model memory.

Validation

  • Complete Go tests passed.
  • go vet passed.
  • Embedded JavaScript syntax and whitespace checks passed.
  • Static Linux amd64 release build passed.
  • Live read-only validation found both 256K ranks loaded, the frontend serving
    the exact 262,272-token context, and Fast mode in_use with qualified HopID
    9/9. The running model was not stopped, restarted, or reconfigured.

See CHANGELOG.md
for the complete change list.