v1.0.0 — Lykuro Native Inference Engine (pre-release)
Pre-release
Pre-release
Release Notes
v1.0.0 (2026-08-11)
First release of the Lykuro Native Inference Engine — a from-scratch
local LLM inference engine with no third-party inference runtime, per
LYK-NIE-SD-001 and the Metal addendum LYK-NIE-ADD-METAL-001.
Highlights
- Own engine, end to end. In-tree JSON parser, SHA-256, safetensors
reader, byte-level BPE tokenizer, chat template, sampler, scheduler,
KV cache, and all inference kernels. No Ollama / llama.cpp / vLLM /
TGI / mlx-lm in source, binary, or transitive form (AT-01 / AT-M03,
enforced bytools/check_no_forbidden_runtime.sh). - Correctness first. CPU FP32 reference as the oracle; every backend
is verified against the real Qwen2.5-0.5B-Instruct checkpoint using an
HF-transformers reference (logits + greedy agreement), and is
bit-exact run-to-run. - Two GPU backends behind one interface.
- CUDA (Linux): BF16-resident weights, fused decode kernels, CUDA
Graphs, paged KV + scoped prefix cache, INT8/INT4 weight-only
quantization, and a 2-way tensor-parallel PoC. ~150 tok/s single /
~450 tok/s at batch 16 on an RTX 3060 (dev figures). - Metal (macOS, Apple Silicon): MPSGraph forward with resident
unified-memory buffers, chunked prefill, graph pre-warm, and FP16
(weights/activations/KV, FP32-safe reductions) halving resident VRAM.
- CUDA (Linux): BF16-resident weights, fused decode kernels, CUDA
- Security & serving. mTLS gRPC with Control/Data identity
separation, Ed25519 manifest signature verification (fail-closed),
content-free logs/metrics, loopback metrics endpoint, and a strict
fail-closed JSON config. - Stability — both 24h soaks passed. CUDA: 237,009 requests, 0 failed,
RSS bit-identical, 100 reload cycles leak-free. Metal: 137,440 completed,
0 failed, phys_footprint flat (~2656MB over 24h) after fixing a
per-request autorelease leak; 100 reload cycles leak 11.7MB. Per-request
bookkeeping is bounded by in-flight work. - Unified Memory admission (Metal). Load-time weight admission plus
runtime staged watermarks — soft sheds new sequences, hard caps KV
growth — with committed-KV accounting. - Supply-chain gate.
tools/scan_vulnerabilities.shcross-checks the
SBOM against a VEX ledger and OSV, fails closed on unreviewed deps or
un-waived CRITICAL/HIGH; emitsvulnerability-report.jsonin CI.
Ollama-style unified commands
Every operation is a native-engine <subcommand>:
pull <hf_repo> [out]— download a HF checkpoint and convert it to a
Lykuro artifact natively (no Python).run <model_or_repo> ["prompt"] [--backend … --max-tokens N --temperature T --system "…"]— config-less inference. Accepts a local artifact dir OR a
HF repo id, which auto-pulls if not already local (single-command run).
Streams output, or starts an interactive chat with no prompt.serve --config <path>— the gRPC engine server (legacy--configstill
works).convert <hf_dir> <out_dir>— HF checkpoint → artifact.serve --http [--port 11434]— Ollama- and OpenAI-compatible HTTP API
(/api/generate, /api/chat, /api/tags, /api/pull; /v1/chat/completions,
/v1/completions, /v1/models). Models given by HF repo id auto-pull.
Packaging
linux-cudaandmacos-metalprofiles viatools/make_package.sh:
staged tree, sorted checksums, provenance manifest, Ed25519-signed
manifest, deterministic tarball. SBOM (SPDX 2.3) and full license
texts included.- Single self-contained binary. gRPC/protobuf/abseil/OpenSSL are
statically linked (release-staticpreset,third_party/build_grpc_static.sh);
macOS links only/usr/lib+ Apple frameworks, Linux the C/C++ runtime- system OpenSSL + CUDA. A cross-platform gate
(tools/check_selfcontained.sh) fails the build on any forbidden dynamic
dependency.
- system OpenSSL + CUDA. A cross-platform gate
- macOS install (§24). pre/postinstall prechecks,
enable_service.sh
(non-root_lykuroLaunchDaemon), anduninstall.sh. Two-phase signing:
Phase 1--devad-hoc (internal test), Phase 2 Developer ID + notarize. - Operations runbook (
docs/operations/runbook.md): deploy, monitor,
Drain/Resume update, versioned-package rollback, recovery.
Known limitations / not certified
- This is a pre-release. Public signed binaries are not attached:
macOS Developer ID signing + Apple notarization (Phase 2,
downloads.lykuro.ai) await Apple Developer Program enrollment. The
sign/notarize/staple pipeline is implemented and self-skipping until the
certificate is present. Build from source with therelease-static
preset, or use the internal Phase 1--devad-hoc package. - Certified Profiles are dev-measured, not production-certified:
formal security review, signed-artifact-only measurement, and
cross-host variance data are pending (the 24h soaks themselves passed). - Custom Metal kernels (precompiled
metallib) require an Xcode CI
runner — the MVP is MPSGraph-only. - Hardware/OS coverage is limited to the entries in
docs/compatibility-matrix.md. Sharded weights, MoE, embeddings,
vision, and NCCL-based multi-GPU are out of scope for this release. - Quantized (INT8/INT4/FP16) models change greedy output relative to the
FP32 oracle by design; per-model quality gating belongs to the offline
evaluation pipeline.
See docs/DEFINITION_OF_DONE.md for the full per-item status against
both specifications.
Quick start
curl -fsSL https://raw.githubusercontent.com/lykuroai/engine/main/deploy/macos/install.sh | bash # macOS, no cert
native-engine pull Qwen/Qwen2.5-1.5B-Instruct
native-engine list
native-engine run Qwen/Qwen2.5-1.5B-Instruct "What is 2+2?"
native-engine serve # HTTP API (Ollama /api/* + OpenAI /v1/*) on 127.0.0.1:11434
native-engine serve --host 0.0.0.0 # expose on the internal LAN (unauthenticated — firewall it)
Downloads
| File | Platform | Backend | SHA-256 |
|---|---|---|---|
lykuro-native-engine-linux-cuda-1.0.0.tar.gz |
Linux x86_64 + NVIDIA | CUDA GPU | 58da75f91e25e800f30062a08849b6b294ece714ee403c76c5d2df214c9f6d3a |
lykuro-native-engine-linux-amd64 |
Linux x86_64 (AMD/Intel) | CPU | 1d05c55c3f076a1a65ba748fd49f815de661a44fa7a497a1d4f04043c7d9ad20 |
lykuro-native-engine-linux-arm64 |
Linux aarch64 | CPU | 620203dc0882d6c0213d086a5fa25e1d82d9a826a6d61d84aac52f2efbdf18e9 |
lykuro-native-engine-macos-arm64 |
macOS Apple Silicon | Metal (Mac GPU) | 239de55524d5ba5031bbaff1afb56200de26d3e68d3a64e1fe2cb0c708828abd |
- macOS ad-hoc signed (not notarized); browser downloads need
xattr -d com.apple.quarantine <file>once, or use the curl installer. - The HTTP API is unauthenticated — bind non-loopback (
--host) only on a
trusted network; use the gRPC mTLS server (serve --config) for auth. - All binaries are single self-contained (static gRPC/protobuf/abseil).