Skip to content

Releases: ciru-ai/CiruStrixLink

CiruStrixLink 0.3.3

Choose a tag to compare

@github-actions github-actions released this 04 Sep 07:29

CiruStrixLink 0.3.3

Version 0.3.3 restores a clean product boundary for the optional Launch page.
CiruStrixLink reports and manages the packaged GLM pair; it does not inspect
unrelated applications or encode a particular operator's service inventory.

Launch behavior

  • Each machine reports its GLM rank, process state, context profile, DFlash
    setting, prefix-cache setting, system-memory use, and configured KV cache.
  • Paired load and unload retain their fixed ordering, readiness checks, bounded
    execution, and rollback behavior.
  • The loader still prevents the portable and NHI forms of this same GLM
    deployment from running together.
  • Other applications and services remain the operator's responsibility and are
    outside CiruStrixLink's policy.

Performance wording

DFlash speed varies with how many proposed tokens the target accepts. The
recorded HumanEval 0–9 gate passed 10/10 at 26.10 weighted generated tokens/s.
The exact 65,680-token prose-heavy recovery probe reached 15.12 tokens/s at
k=5, versus 12.16 at k=3 and 9.38 target-only. That recovery request had much
lower draft acceptance and is documented as a stress case, not as the model's
normal decode rate.

Runtime scope

This release changes no model weights, vLLM code, DFlash implementation,
kernel, launch recipe, context profile, KV allocation, or USB4 transport.

Validation

  • Complete Go test suite
  • go vet
  • Embedded JavaScript syntax validation
  • Whitespace checks
  • Static Linux amd64 release build

The already-running GLM pair was not stopped, restarted, reconfigured, or sent
validation traffic while this correction was prepared.

CiruStrixLink 0.3.2 — superseded

Choose a tag to compare

@github-actions github-actions released this 04 Sep 03:52

CiruStrixLink 0.3.2 — superseded

Version 0.3.2 added host-specific workload detection that did not belong in a
general CiruStrixLink release. Version 0.3.3 removes that behavior while keeping
the paired Launch workflow and two-rank runtime reporting intact.

Users should skip 0.3.2 and install 0.3.3 or newer.

v0.3.1

Choose a tag to compare

@github-actions github-actions released this 04 Sep 03:27

CiruStrixLink 0.3.1

  • Replaced hardcoded DFlash and prefix-cache labels with settings reported by
    both TP ranks.
  • Reported Fast USB4 state separately from DFlash state.
  • Recorded the 256K DFlash2 k=5, prefix-cache-off serving configuration.

Benchmark methodology note

One long-context DFlash stress campaign was later found not to have full host
isolation. Its throughput comparison is directional evidence rather than an
isolated-system measurement. The HumanEval 0–9 correctness result remains
10/10. This note does not represent a runtime or model change.

CiruStrixLink 0.3.0

Choose a tag to compare

@ciru-ai ciru-ai released this 03 Sep 19:21

CiruStrixLink 0.3.0

Version 0.3.0 turns the embedded console into a practical two-machine operator
surface while preserving CiruStrixLink's fail-closed transport behavior.

Highlights

  • Redesigned Connection Setup as a guided two-host workflow.
  • Replaced Launch Settings with a real Launch page and moved it to the third
    tab.
  • Added coordinated configure, load, unload, readiness, and rollback handling
    for GLM5.3 Flash CIRU STRIX IU4.
  • Added live rank, process, context, KV allocation, host-wide unified RAM, and
    available-memory telemetry for both machines.
  • Fixed Fast-mode reporting so Overview, Diagnostics, and Launch use the same
    scoped privileged truth when model control is enabled.
  • Added copyable permission setup for packaged NixOS and generic Linux.
  • Added a built-in, root-only generic Linux helper with exact sudoers commands.

Safe paired lifecycle

Model control is opt-in and requires a shared peer token, complementary fixed
ranks, fixed USB4 IPv4 peers, a model frontend URL, a loopback-only console,
and narrow root-owned helpers on both hosts.

A load verifies both ranks and the NHI lease, writes the same selected profile
to both stopped ranks, starts rank 0 followed by rank 1, and waits until the
frontend serves the exact context. Failure stops both ranks again and reports
any incomplete rollback. Unload stops rank 1 before rank 0.

CiruStrixLink refuses to load this deployment while qwen-main.service or the
portable GLM service is active. It never enables, disables, masks, or replaces
those services, and it never changes a context while the model is running.

Desktop access

Run the console on one model host and the USB4-bound agent on the other. From a
desktop, forward the console's loopback port:

ssh -N -L 7749:127.0.0.1:7749 USER@CONSOLE_MODEL_HOST

Open http://127.0.0.1:7749/#/launch. The desktop needs no StrixLink binary or
model-host privileges.

See the browser-console instructions,
GLM deployment guide,
and security model.

256K profile

The released 256K profile configures 262,272 tokens with an 8 GiB KV allocation
per rank and 8 GiB of host prefix staging per rank. It is a single-request,
experimental high-context profile. The UI reports host-wide unified memory so
operators can see the actual remaining headroom rather than treating the KV
allocation as total model memory.

Validation

  • Complete Go tests passed.
  • go vet passed.
  • Embedded JavaScript syntax and whitespace checks passed.
  • Static Linux amd64 release build passed.
  • Live read-only validation found both 256K ranks loaded, the frontend serving
    the exact 262,272-token context, and Fast mode in_use with qualified HopID
    9/9. The running model was not stopped, restarted, or reconfigured.

See CHANGELOG.md
for the complete change list.