Distributed inference fabric for custom LLM systems. Route workloads across CPU, GPU, and NPU resources with explicit scheduling, hardware-aware placement, and a self-hosted open-source stack.
Quick Start · Demo · Architecture · API · Docs · Contributing
Install → load a model → chat, end to end, running Llama 3.2 1B locally through Ghostlink Studio's native llama.cpp backend.
The full-length recording (with audio) lives in the gitignored
demo/folder for local viewing — it's not part of the pushed repo, so there's no link to it here.
- What Ghostlink brings
- Why Ghostlink
- Quick Start
- Hardware Detection & Compatibility
- Comparison vs. Other Platforms
- Performance
- Launch Scripts
- Usage (Developer)
- MCP Tools
- Editor Tab & Copilot Features
- Troubleshooting
- Environment Variables
- Architecture
- Testing
- Project Status
- Contributing
- License
- Clear routing and scheduling for custom inference topologies
- Real cross-machine distributed inference: opt in a node's spare GPU/CPU and
a request automatically splits across it via llama.cpp's own RPC backend —
zero-config, no manual
--rpcflags - Hardware-aware placement across mixed compute environments (CPU, GPU, NPU)
- SPSC ring-buffer transport with spin-wait for sub-microsecond handoff
- TCP and Unix domain socket transport for multi-process pipelines
- Session-level authentication on transport frames
- Dynamic system profile watching with auto-tuning cache
- In-GUI code editor (Monaco) with copilot-style actions — Explain/Fix/Refactor with diff preview before anything is written, multi-file refactor, opt-in ghost-text autocomplete, and repo-aware chat via local RAG indexing
- A strong open-source foundation with a commercial support path
Point Ghostlink at every machine on your LAN and it becomes one inference
cluster — correctly sized, authenticated, and observable — with zero YAML
and no manual --rpc flags. Nobody else combines zero-config discovery of
heterogeneous consumer/prosumer hardware (gaming GPU + old laptop +
NPU-equipped ultrabook + Mac) with real distributed inference across it:
vLLM assumes a homogeneous co-located GPU fleet, Ollama and LM Studio don't
distribute at all, and Kubernetes-based serving solves this but needs a
cluster and an ops team. See docs/COMPARISON.md for
the full breakdown and docs/BENCHMARKS.md for a real
two-machine LAN run backing this up, not just prose.
Use Ghostlink when you need:
- zero-config cluster formation across mixed CPU/GPU/NPU hardware already sitting on your network,
- a self-hosted inference fabric with open-source flexibility and no vendor lock-in,
- a platform that can be extended into a paid commercial offering with support and enterprise deployment services.
Three commands, one launcher, no manual service wiring — launch.bat / launch.sh builds what's missing, starts every process, and opens the GUI.
No Rust/Node/cmake toolchain needed — this downloads the prebuilt ghost-link
binary from GitHub Releases,
verifies it against the release's published SHA256SUMS, and installs it to
a user-writable directory (no sudo/admin required).
Linux / macOS:
curl -fsSL https://raw.githubusercontent.com/rwilliamspbg-ops/Ghostlink/main/scripts/install.sh | shWindows (PowerShell):
irm https://raw.githubusercontent.com/rwilliamspbg-ops/Ghostlink/main/scripts/install.ps1 | iexThen:
ghost-link --help
ghost-link doctor --strict # sanity-check your setup
ghost-link serve 127.0.0.1 8003 # start the OpenAI-compatible API serverBoth scripts accept VERSION=v1.17.0 (env var) to pin a specific release
instead of latest, and print PATH guidance if the install directory isn't
already on it. Source for both lives at
scripts/install.sh and
scripts/install.ps1 — read them before piping to a
shell if you'd rather not blind-trust a one-liner, that's a reasonable
instinct.
What this does not give you: the Go control-plane gateway or the
built React GUI aren't published as standalone release assets today — only
the ghost-link binary, its checksums, and an SBOM are (see
scripts/release_bundle.sh). The one-line
install gets you the CLI and OpenAI-compatible API server; for the full
browser GUI, use the build-from-source or Docker paths below. Release
binaries are also x86_64/amd64-only — there's no arm64 build yet, and
both scripts detect and refuse non-x86_64 hosts rather than silently
handing you a binary that won't run.
Needed for the browser GUI, local development, or platforms without a published release binary.
Prerequisites
| Tool | Required | How to Install |
|---|---|---|
| Rust | Yes | winget install Rustlang.Rustup or https://rustup.rs |
| Node.js | Yes | winget install OpenJS.NodeJS.LTS or https://nodejs.org (LTS) |
| CMake | For llama.cpp | winget install Kitware.CMake or https://cmake.org/download/ |
| Git | For llama.cpp | winget install Git.Git |
Clone and launch
git clone https://github.com/rwilliamspbg-ops/Ghostlink.git
cd Ghostlink
# One-time backend build (if target\release\ghost-link.exe is missing, launch.bat builds it)
cargo build --release -p ghost-link
# Single launcher for the full stack
.\launch.batThis starts:
- llama-server (inference, port 8080)
- Ghostlink API (chat / models / settings, port 8003) — internal; the GUI does not call this directly
- Control-plane (Go gateway: CORS, request logging, rate limiting, streaming-safe proxy, port 8000) — GUI must use this port
- React frontend at http://127.0.0.1:5173, opened automatically in your browser
Verify it's actually running (skip if the browser tab already loaded cleanly):
curl http://127.0.0.1:8000/health # control-plane — GUI-facing gateway
curl http://127.0.0.1:8003/health # ghost-link API — should agree with the line aboveBoth should return a JSON body with a healthy status, not a connection error.
Open http://127.0.0.1:5173 → Models → load a model (a small one like Llama-3.2-1B-Instruct loads in seconds and is a good first check) → Chat.
Optional:
# Ollama instead of llama-server
$env:GHOSTLINK_INFERENCE_BACKEND="ollama"; .\launch.bat
# Fall back to the old WSL-delegated launcher (runs launch.sh inside WSL)
$env:GHOSTLINK_USE_WSL="1"; .\launch.bat405 on chat/models? The GUI was pointed at the wrong port (e.g. ghost-link
:8003directly, or llama-server:8080). Always use API basehttp://127.0.0.1:8000(the control-plane gateway).Chat suddenly errors with "error sending request for url (...8080...)"?
llama-serverdied — usually from two model-load requests overlapping (double-clicking Load, switching models fast). Fixed in 1.7.1+; if you're still seeing it,Unloadthen re-Loadthe model from the Models tab.
Prerequisites: Rust, Node.js, curl; cmake optional (a prebuilt llama-server is used as a fallback if cmake isn't available).
git clone https://github.com/rwilliamspbg-ops/Ghostlink.git
cd Ghostlink
cargo build --release -p ghost-link
./launch.sh
# launch-complete.sh is a thin wrapper around launch.shSame verification and port layout as Windows above — curl http://127.0.0.1:8000/health, then http://127.0.0.1:5173 → Models → Chat.
No Rust/Node/cmake toolchain needed — just Docker.
git clone https://github.com/rwilliamspbg-ops/Ghostlink.git
cd Ghostlink
mkdir -p models
curl -L -o models/stories15M-q4_0.gguf https://huggingface.co/ggml-org/models/resolve/main/tinyllamas/stories15M-q4_0.gguf
docker compose up --buildThen open http://localhost:5173 → Models → Chat. Swap models/stories15M-q4_0.gguf (and the llama-server volume mount in docker-compose.yml) for a larger GGUF once you've confirmed the stack comes up.
docker-compose.demo.ymlis an older prototype (Python model manager + Ollama) and is no longer maintained — usedocker-compose.ymlabove.
Ghostlink auto-detects available accelerators at startup by probing in this order: nvidia-smi (cross-platform), then an OS-specific fallback (Windows WMI / Linux lspci+sysfs / macOS system_profiler), then a Vulkan probe as a last resort. What's actually detectable, per OS (grounded in system_profile.rs):
| Accelerator | Windows | Linux | macOS | Detection |
|---|---|---|---|---|
| CUDA (NVIDIA) | ✅ | ✅ | nvidia-smi, CUDA_PATH |
|
| DirectML (AMD iGPU, Intel ARC, any DX12 GPU) | ✅ | ❌ | ❌ | WMI |
| ROCm (AMD discrete) | ❌ | ✅² | ❌ | rocm-smi, hipconfig — needs --features rocm |
| Vulkan (generic GPU fallback) | ✅ | ✅ | ✅ | Vulkan probe, full-scan mode only |
| Metal (Apple Silicon) | ❌ | ❌ | ✅ | sysctl hw.optional.arm64 |
| NPU (AMD XDNA, Intel NPU, Qualcomm) | ✅ | ✅ | ✅³ | WMI (Windows), sysfs//sys/class/accel (Linux), Apple Neural Engine (macOS) |
| AF_XDP kernel bypass | ❌ | ✅ | ❌ | Linux-only transport optimization (xdp.rs) |
| CPU | ✅ | ✅ | ✅ | Always available — the guaranteed fallback |
¹ NVIDIA GPU support on macOS has been effectively deprecated by both Apple and NVIDIA for years — the probe runs, but don't expect it to find anything on real hardware.
² ROCm detection is compiled out entirely unless you build with cargo build --features rocm; without it, AMD discrete GPUs on Linux fall through to the Vulkan probe.
³ Detected as the Apple Neural Engine, not a discrete accelerator — no VRAM/memory figure is reported since it shares system memory.
Backend support (native llama-server vs. Ollama): both work identically on all three OSes — Ollama just needs the ollama binary on PATH (GHOSTLINK_INFERENCE_BACKEND=ollama), no build changes required.
If your GPU isn't detected, set env vars manually:
$env:GHOSTLINK_GPU_NAME="AMD Radeon 860M"
$env:GHOSTLINK_VRAM_GB=8
$env:GHOSTLINK_COMPUTE_CAPABILITY="gpu"A full hardware probe (spawning several external commands: PowerShell/CIM
queries on Windows, nvidia-smi, etc.) runs concurrently rather than
sequentially — measured at ~1.4s on a Windows dev machine, down from ~4s when
these ran one after another with no shared state between them.
Sub-microsecond core primitives drive Ghostlink's distributed inference fabric. Benchmarks
run on an Intel i7-14700K (Linux/WSL2) with RUSTFLAGS="-C target-cpu=native".
| Benchmark | Latency | Throughput |
|---|---|---|
| Ring buffer push+pop (single-thread) | 1.94 ns | 516 M ops/s |
| Ring buffer SPSC cross-thread (100k) | 10.67 ns | 93.7 M ops/s |
| Protocol: DiscoveryFrame encode | 75.80 ns | 13.2 M ops/s |
| Protocol: DiscoveryFrame decode | 126.82 ns | 7.9 M ops/s |
| Protocol: encode + decode round-trip | 196.46 ns | 5.1 M ops/s |
| Planning: 33 layers / 2 nodes | 115.46 ns | 8.7 M ops/s |
| Planning: 80 layers / 8 nodes | 245.31 ns | 4.1 M ops/s |
| Planning: autotuned (80 layers, 8 nodes) | 360.99 ns | 2.8 M ops/s |
| Cluster: register node | 193.04 ns | 5.2 M ops/s |
| Cluster: snapshot (10 nodes) | 533.40 ns | 1.9 M ops/s |
| Autotune: detect runtime profile | 166.58 ns | 6.0 M ops/s |
| Transport | Tokens / batch | Throughput | Latency |
|---|---|---|---|
| In-process (spin-wait) | 1024 / 128 | 900 K tok/s | 1.14 ms |
| TCP loopback | 1024 / 128 | 340 K tok/s | 3.01 ms |
| In-process (spin-wait) | 256 / 32 | 639 K tok/s | 0.40 ms |
| TCP loopback | 256 / 32 | 236 K tok/s | 1.08 ms |
These four numbers move a 64-byte/token synthetic payload — a stand-in roughly 128x smaller than a real model's per-token activation, so treat this table as a transport-layer ceiling, not an LLM-workload number. For realistic payload sizes (4K-32K token batches at real FP16/BF16 activation byte sizes, P99 latency, bandwidth GB/s, and a Ray actor-transfer comparison baseline), see docs/BENCHMARKS.md's "LLM-Shaped Workload Benchmarks" section.
cargo bench -p ghostlink-core --bench criterion # microbenchmarks
python scripts/flow_perf_snapshot.py --exec-tokens 512 --micro-batch 8 --runs 5 --modes tcp inmem --releaseNote: the flow command's TCP/XDP paths can now carry a realistic
per-token payload — opt-in via GHOSTLINK_FLOW_HIDDEN_DIM/GHOSTLINK_FLOW_DTYPE_BYTES
(e.g. 4096/2 for a 7B-class FP16/BF16 model), not the default. The
command above still reproduces the legacy Pipeline Throughput table's
64-byte/token numbers unchanged; see the LLM-Shaped section linked above for
the exact reproduce command with those env vars set.
Every number above is falsifiable — run the commands yourself. Numbers vary meaningfully by hardware: docs/BENCHMARKS.md documents two more full profiles with the same rigor — an AMD Radeon 860M integrated GPU laptop (Windows) and a 4-core Linux mini PC with no dedicated GPU — including honest noisiness notes for both low-power classes. A first real multi-node LAN run (that same Windows laptop and Linux mini PC, over a real network) is also in BENCHMARKS.md's Multi-Node Performance section; more node counts and hardware pairs (especially discrete GPUs) are still open, see docs/ROADMAP.md, Horizon 1.
Native nodes enable Flash Attention, VRAM-scaled batch sizes, and Q8_0 KV cache by default. Prefer Q4_K_M / IQ4_XS over FP16/Q8_0 for ~1.5–2× decode speed. Override with GHOSTLINK_LLAMA_SERVER_ARGS. See docs/LOCAL_INFERENCE_TUNING.md.
GPU layer offload (-ngl) is now always passed explicitly to llama-server,
including the "auto-detect" case (-ngl -1, let llama-server decide) when no
VRAM/GPU is detected or the corresponding env vars aren't set. Previously the
flag was omitted entirely in that case, and llama-server's own default
(-ngl absent) is CPU-only — meaning inference silently ran on CPU with no
GPU offload and no warning on any launch path that didn't set
GHOSTLINK_VRAM_GB/GHOSTLINK_LLAMA_NGL itself.
That auto-detect default is itself capped to CPU-only for large (≥10GB)
models on an integrated GPU, regardless of VRAM tier — measured directly,
full offload on this class of hardware duplicates weights into RAM rather
than freeing it the way a discrete GPU would, and can leave a host with
under 1GB free while a large model is loaded. GHOSTLINK_LLAMA_NGL always
overrides this either way. See
docs/LOCAL_INFERENCE_TUNING.md for the
measured numbers and the reasoning.
GPU offload, context length, threads, batch/µ-batch size, KV cache quantization, Flash Attention, and mlock/no-mmap can all be tuned live from the GUI's Settings tab ("Model Performance" section) as well as via env vars — both paths default to "Auto" for every value that already has a measured, VRAM/model-size-aware default, so touching the GUI never silently disables the safety caps above unless you explicitly opt out of "Auto" for that field. See docs/LOCAL_INFERENCE_TUNING.md for the full env var list and a measured context-length-vs-full-offload ceiling on integrated GPUs.
Compiling with RUSTFLAGS="-C target-cpu=native" further improves performance by enabling CPU-specific instruction sets (opt-in; not set by default — a multi-node cluster can't assume every node shares one CPU microarchitecture).
Release builds use lto = "thin" and codegen-units = 1 ([profile.release] in the workspace Cargo.toml) so the compiler can inline across the ghostlink-core / ghost-link crate boundary — the ring buffer, protocol, and planning hot paths all cross it. A same-machine, same-run A/B (not the table above, which is a different machine/OS) showed the deterministic single-threaded paths 6–19% faster with this profile; thread-scheduling-bound benchmarks were noisier and not clearly attributable to it either way. panic = "abort" was considered and deliberately not set, since it would crash the whole server process on any panic instead of failing just one request.
| Script | Description |
|---|---|
launch.bat |
Windows — native launcher (llama-server + API :8003 + control-plane gateway :8000 + GUI :5173, no WSL). Set GHOSTLINK_USE_WSL=1 to use launch.sh inside WSL instead. |
launch-native.ps1 |
The actual native-Windows implementation launch.bat calls into |
launch.sh |
Linux/macOS — same stack with hardware detection |
launch-complete.bat / launch-complete.sh |
Compatibility wrappers → launch.bat / launch.sh |
launch-ollama.bat |
Thin wrapper setting GHOSTLINK_INFERENCE_BACKEND=native (both bat and ps1 launchers respect it; set GHOSTLINK_INFERENCE_BACKEND=ollama yourself for Ollama) |
# Build
cargo build --release -p ghost-link
# With AMD ROCm support
cargo build --release -p ghost-link --features rocm
# Probe local hardware profile (with auto-tuning cache)
cargo run -p ghost-link -- probe my-node
cargo run -p ghost-link -- probe my-node --full
# Generate a placement plan for your hardware
cargo run -p ghost-link -- plan
# Start the OpenAI-compatible API server
cargo run -p ghost-link -- serve 127.0.0.1 8003
# Unified troubleshooting
cargo run -p ghost-link -- doctor --strict
# Run the full 30B planning flow
cargo run -p ghost-link -- flow iprada-16gb zenbook-32gb 32 32 64 4 tcp
# Launch the ASCII cluster dashboard
cargo run -p ghost-link -- dashboardEvery route below except /health requires Authorization: Bearer <token> —
either the API key printed once at server startup (also saved to
api_key.txt), or a short-lived JWT exchanged for it via
POST /api/security/jwt/refresh. See docs/API_REFERENCE.md
for full request/response examples and docs/openapi.yaml
for a machine-readable spec.
| Endpoint | Method | Description |
|---|---|---|
/health |
GET | Health check (no auth required) |
/v1/chat/completions |
POST | OpenAI-compatible chat completion |
/v1/completions |
POST | OpenAI-compatible legacy (prompt-based) completion |
/v1/embeddings |
POST | OpenAI-compatible embeddings (Ollama backend only) |
/v1/models |
GET | OpenAI-compatible model list |
/api/models |
GET | List available models |
/api/models/status |
GET | Loaded model status |
/api/runtime/detect |
GET | Available runtimes (GPU, NPU, CPU) |
/api/runtime/models?runtime=directml |
GET | Models filtered by runtime |
/api/backends |
GET | Current backend + available backends |
/api/backends/switch |
POST | Switch inference backend |
/api/backends/:name/status |
GET | Status for one specific backend |
/api/metrics |
GET | Performance metrics |
/api/inference/chat |
POST | Chat completion |
/api/models/partial |
GET | List interrupted downloads (.gguf.part files) |
/api/models/partial/discard |
POST | Delete an interrupted download |
/api/security/jwt/refresh |
POST | Exchange the API key for a short-lived JWT |
/api/security/pqc/state |
GET | Whether this running server is serving HTTPS/PQC-hybrid TLS |
/api/security/pqc/enable |
POST | Persist enable_tls: true (takes effect on next restart) |
/api/security/audit-log |
GET | Security audit log entries (failed auth, JWT refresh, PQC enable, tool-call approve/deny — last 500, most recent first) |
/api/workspace/tree |
GET | List a directory under the Editor tab's workspace root (?path=) |
/api/workspace/file |
GET / PUT | Read or write a workspace file (?path= / {path, content}) — both reject any path that escapes the workspace root |
/api/workspace/index |
POST | Index the workspace into the rag MCP server for repo-aware chat context; "skipped" if rag/Ollama isn't reachable |
/metrics |
GET | Prometheus-exposition-format metrics (same data as /api/metrics) |
A Python client package lives at sdks/python — wraps the OpenAI-compatible endpoints plus native token-streaming chat, workers, metrics, and settings. See sdks/python/README.md.
from ghostlink_client import GhostlinkClient
client = GhostlinkClient("http://127.0.0.1:8003", api_key="<your api key>")
resp = client.chat.completions.create(model="llama3.2:3b", messages=[{"role": "user", "content": "hi"}])
print(resp.content)Ghostlink chat can call real tools via MCP servers,
configured in mcp_servers.toml (a gitignored, per-install copy of
mcp_servers.example.toml, auto-created on first run — same pattern as
ghostlink.toml/ghostlink.example.toml).
Default servers:
| Chat tool slot | Backing server | Enabled by default |
|---|---|---|
file_operations |
@modelcontextprotocol/server-filesystem (npx) |
✅ |
api_call |
mcp-server-fetch (uvx) |
✅ |
calculator |
mcp-calculator (this repo, evalexpr-backed) |
✅ |
database_query |
mcp-server-sqlite (uvx) |
✅ |
web_search |
@modelcontextprotocol/server-brave-search (npx) |
needs BRAVE_API_KEY |
code_execution / terminal |
Docker MCP Toolkit gateway | needs Docker Desktop running |
image_generation |
(not yet configured — no default backend picked) | ❌ |
| — | sequential-thinking (npx) |
✅ |
| — | vision (this repo, wraps local Ollama) |
needs a pulled vision model |
| — | rag (this repo — index_document/search, local Ollama embeddings + brute-force cosine index, no external vector DB) |
✅ (needs a pulled embedding model, e.g. nomic-embed-text — the Editor tab's auto-index degrades to a no-op "skipped" if Ollama isn't reachable) |
The model decides whether and which tool to call (a ReAct-style prompt works
with any local GGUF/Ollama model; Ollama models whose chat template declares
native tool-calling support use that automatically instead). Tools marked
requires_confirmation in mcp_servers.toml (terminal, code_execution) pause
for explicit user approval before executing — see the MCP tab in the GUI.
Requires npx/node (bundled MCP servers) and uvx/python (Python-distributed
ones) on PATH; Docker Desktop for the terminal/code_execution slots.
Ghostlink Studio's Editor tab is a Monaco-based code editor over the
running server's real filesystem (/api/workspace/*, confined to a
canonicalized workspace root — no path traversal outside it), separate from
the sandboxed file_operations MCP tool above.
- Browse, open, and save any file under the workspace root, with syntax highlighting per extension.
- Explain / Fix / Refactor — scoped to the current selection, or the
whole file if nothing's selected. Fix/Refactor render their proposed change
as a side-by-side diff (Monaco's
DiffEditor) with explicit Accept/Reject — nothing is written to disk until you accept it. - Multi-file refactor — select several files in the tree, send them in one prompt, then step through each proposed change individually (Accept/Reject/Skip).
- Ghost-text autocomplete (opt-in, via the lightning-bolt toggle) — Monaco's native inline-completion UI, driven by the same chat-completion endpoint as everything else here. Not true fill-in-the-middle (no suffix awareness, no model-specific FIM tokens) — a real network+inference round trip on a debounce, not per-keystroke-fast.
- Repo-aware chat context — on first load, the Editor tab feeds the
workspace into the
ragMCP server (enabled by default; needs a pulled Ollama embedding model) so chat can pull in relevant file content without anyone callingindex_documentby hand. Re-indexing a file replaces its prior chunks rather than duplicating them.
If Ollama logs show POST /api/generate returning 404:
- Check installed tags:
ollama list - Verify tags through API:
curl http://127.0.0.1:11434/api/tags - Pull the exact tag:
ollama pull qwen2.5:3b - Select that exact tag in Ghostlink before sending chat requests.
- If logs show device visibility overrides:
setx HSA_OVERRIDE_GFX_VERSION "" - Restart Ollama and Ghostlink after model or environment changes.
Launch scripts check for port conflicts before binding. If you see "address already in use", ensure no stale processes are holding ports 8000, 8003, or 8080.
| Variable | Default | Description |
|---|---|---|
GHOSTLINK_INFERENCE_BACKEND |
native |
native or ollama |
GHOSTLINK_NATIVE_ENGINE |
llama_server |
llama_server or llama_cpp |
GHOSTLINK_LLAMA_SERVER_URL |
http://127.0.0.1:8080/completion |
llama-server URL |
GHOSTLINK_GPU_NAME |
— | Override detected GPU name |
GHOSTLINK_VRAM_GB |
— | Override detected VRAM |
GHOSTLINK_COMPUTE_CAPABILITY |
— | Override detected compute capability |
GHOSTLINK_RUNTIME |
— | Force a runtime selection |
GHOSTLINK_FORCE_RUNTIME |
false |
When true, honor GHOSTLINK_RUNTIME even if not auto-detected |
GHOSTLINK_SYSTEM_MEMORY_GB |
— | Override detected system memory |
NPU_DEVICE / QUALCOMM_NPU |
— | Enable NPU detection via env |
See ghostlink.toml for all settings:
- Node identities and resource overrides
- Discovery broadcast configuration
- TCP transport tuning (max_inflight, auth_token, reconnect)
- GUI Python path
See docs/ARCHITECTURE.md for a request/cluster-flow diagram.
crates/
├── ghostlink-core/ # Shared runtime primitives
│ ├── ring.rs # SPSC lock-free ring buffer (spin-wait)
│ ├── runtime.rs # Pipeline execution (in-memory, TCP, Unix, AF_XDP)
│ ├── system_profile.rs # Cross-platform GPU/NPU/CPU auto-detection
│ ├── autotune.rs # Auto-tuning cache with hardware fingerprinting
│ ├── watcher.rs # Dynamic hot-plug system profile watcher
│ ├── host.rs # Compute host summary
│ ├── planning.rs # Layer assignment & quantization
│ ├── protocol.rs # Binary frame protocol with CRC32 + HMAC auth
│ ├── discovery.rs # UDP broadcast cluster discovery
│ ├── mdns.rs # mDNS cluster discovery (VLAN/VPC-friendly)
│ ├── cluster.rs # Thread-safe node state & metrics
│ ├── health.rs # Network health & fault detection
│ ├── load_balance.rs # Tensor distribution across nodes
│ ├── accelerator.rs # NPU acceleration detection
│ └── xdp.rs # AF_XDP kernel bypass (fallback-safe)
├── ghost-link/ # CLI demo & API server
│ └── src/mcp/ # MCP client (rmcp): registry, config, tool-calling loop
├── mcp-calculator/ # Custom MCP server backing the calculator chat tool
├── mcp-vision/ # Custom MCP server backing the vision chat tool (local Ollama)
ghostlink_gui_modern/ # React frontend (Vite + Tailwind)
# Full test suite
cargo test --workspace
# With ROCm feature
cargo test --workspace --features rocm
# Run benchmarks
cargo bench --package ghostlink-coreBefore pushing to CI, run:
cargo fmt --all --check
cargo clippy --workspace --all-targets -- -D warnings
cargo test --workspaceCI enforces the same checks across Ubuntu, Windows, and macOS.
Ghostlink is positioned as a launch-ready open-source foundation with a strong demo story and public-facing collateral.
Current strengths:
- a working local launch path for experimentation and demos,
- a clear positioning around distributed inference scheduling and routing,
- public assets for comparison, demo flow, and product storytelling.
Current focus areas:
- strengthening the end-to-end demo experience,
- improving documentation for deployment and production use,
- expanding real-world validation across more hardware and runtime setups.
Public launch assets:
- Live site: https://rwilliamspbg-ops.github.io/Ghostlink/
- Comparison sheet: docs/COMPARISON.md
- Demo flow: docs/archive/launch_demo.md
See CONTRIBUTING.md for setup, PR expectations, and release rubric. See CONTRIBUTORS.md to see who's built this so far.
MIT
