From 7b5a7806b06b7e446f203819fc8d580aac031949 Mon Sep 17 00:00:00 2001 From: M2Night Date: Fri, 7 Aug 2026 00:12:56 +0800 Subject: [PATCH] Add enterprise appliance and release-history pages MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Self-host enterprise customers had nowhere public to read how to run the appliance or what changed between versions. The existing Self-Hosting section covers the open-source fish-speech images, which is a different product for a different audience. Two pages: - Enterprise Appliance — prerequisites, sign in and pull (including the air-gapped save/load path), the run command with the reasons for each flag that matters, and upgrading. - Appliance Releases — the version list. The dashboard links here from "a newer version is available", so each entry has to say what changed and what an upgrade involves, not just a date. Content comes from the operator guide we have been sending customers by hand, trimmed to what a customer needs: no registry internals, no build pipeline. Two things deliberately left out because they are commitments only the business can make: a support window (how long an old version keeps getting fixes) and any promise that older tags stay pullable indefinitely. Both should be added once decided — the releases page currently points people at their account manager. The version list has one entry. It needs a real second entry the first time we ship an upgrade, and the "how to read a version" section assumes the naming scheme stays as it is. Co-Authored-By: Claude Opus 5 --- .../self-hosting/enterprise-appliance.mdx | 122 ++++++++++++++++++ .../self-hosting/enterprise-releases.mdx | 49 +++++++ docs.json | 4 +- 3 files changed, 174 insertions(+), 1 deletion(-) create mode 100644 developer-guide/self-hosting/enterprise-appliance.mdx create mode 100644 developer-guide/self-hosting/enterprise-releases.mdx diff --git a/developer-guide/self-hosting/enterprise-appliance.mdx b/developer-guide/self-hosting/enterprise-appliance.mdx new file mode 100644 index 0000000..4ab163f --- /dev/null +++ b/developer-guide/self-hosting/enterprise-appliance.mdx @@ -0,0 +1,122 @@ +--- +title: "Enterprise Appliance" +description: "Run the Fish Audio TTS stack on your own hardware from a single container" +icon: "server" +--- + +The enterprise appliance is the whole Fish Audio TTS stack — model weights included — in one +container. It needs no Kubernetes and no network access at runtime, which makes it suitable for +on-prem, single-tenant, and air-gapped deployments. + + + The appliance is part of an enterprise agreement. Your team needs the **Self + Host** feature and a grant for the All-in-One artifact before the commands + below will work. If Developer → Self Host does not appear in your dashboard, + contact your account manager. + + +## What you get + +| | | +| ------------------ | -------------------------------------------------------------------------------------------------------------------------------------------------------- | +| Exposed port | `8088` — the TTS API. Nothing else is published. | +| GPUs | **2**, validated on 2× RTX 5090 (32 GB) and 2× H100 80 GB. GPU 0 runs the language model, GPU 1 runs the vocoder. No NVLink required. | +| Scaling | one inference worker on two GPUs. It does not shard across more GPUs or nodes — for elastic or multi-tenant throughput, use the Helm deployment instead. | +| Network at runtime | none. Weights are baked into the image. | + +## 1. Prerequisites + +- Docker with the **NVIDIA Container Toolkit** installed and the `nvidia` runtime registered. + Verify with `docker run --rm --gpus all nvidia-smi`. +- Enough disk for the image: roughly **28 GB compressed, 60 GB unpacked**. +- A **deploy token**, created in Developer → Self Host. That page also shows the version to run. + +## 2. Sign in and pull + +Your username is your Fish Audio account email; the password is a deploy token. + +```bash +echo '' | docker login registry.fish.audio -u '' --password-stdin +docker pull registry.fish.audio/self-hosted/enterprise/all-in-one: +``` + + + Developer → Self Host renders both of these commands with your email and the + current version already filled in. Copy them from there rather than typing the + version by hand. + + +For an air-gapped host, pull on a machine that can reach the registry, then move the image: + +```bash +docker save registry.fish.audio/self-hosted/enterprise/all-in-one: \ + | zstd -T0 -3 -o all-in-one.tar.zst +# on the target host: +zstd -d -c all-in-one.tar.zst | docker load +``` + +## 3. Run + +Generate a JWT secret **once**, store it, and reuse the same value on every run. + +```bash +export FISH_JWT_SECRET="$(openssl rand -hex 32)" + +docker run -d --name fish-tts \ + --gpus all \ + --shm-size 16g --ulimit memlock=-1 --ulimit stack=67108864 \ + -p 8088:8088 \ + -v fish-tts-shared:/mnt/shared \ + -e JWT_SECRET="$FISH_JWT_SECRET" \ + --restart unless-stopped \ + registry.fish.audio/self-hosted/enterprise/all-in-one: +``` + + + `JWT_SECRET` is required for any production deployment. Without it the + container falls back to a fixed built-in development default, which is not + secret. Changing the value later invalidates every token and session issued + under the old one. + + +**`--gpus all`** pins the worker to GPU 0 and the vocoder to GPU 1. On a host with more than two +GPUs it takes the first two; to choose specific cards use `--gpus '"device=0,1"'`. + +**`-v fish-tts-shared:/mnt/shared`** is one persistent volume for everything that must survive a +restart: the compile and CUDA-graph caches, the vocoder engine, reference voices, and the usage +ledger. It is what makes restarts fast. + + + The **first start compiles for around 10 minutes** once the image is on the + host. A completely cold host that also has to transfer the ~28 GB image can + take 45–75 minutes end to end, depending on the network. Later starts on the + same volume take minutes. Keep the volume. + + +The container runs fully non-root — PID 1 and every service as UID 1000. A fresh named volume +inherits that ownership and works as-is; an existing volume or a host bind-mount must be writable +by UID 1000. + +## 4. Check it is serving + +```bash +curl -fsS http://localhost:8088/health +``` + +## Upgrading + +Developer → Self Host shows the version your team recorded and tells you when a newer one is +available, with a link to what changed. Upgrading is: pull the new tag, stop the old container, +start a new one with the same volume and the same `JWT_SECRET`. + +```bash +docker pull registry.fish.audio/self-hosted/enterprise/all-in-one: +docker stop fish-tts && docker rm fish-tts +# re-run the command from step 3 with the new tag +``` + +The volume is reused deliberately: the caches in it are keyed by content, so a new image rebuilds +only what actually changed. Keep the old image on the host until the new one has served traffic — +rolling back is then just starting the previous tag again. + +See [Appliance Releases](/developer-guide/self-hosting/enterprise-releases) for the version list. diff --git a/developer-guide/self-hosting/enterprise-releases.mdx b/developer-guide/self-hosting/enterprise-releases.mdx new file mode 100644 index 0000000..0285442 --- /dev/null +++ b/developer-guide/self-hosting/enterprise-releases.mdx @@ -0,0 +1,49 @@ +--- +title: "Appliance Releases" +description: "Version history for the Fish Audio enterprise appliance" +icon: "tag" +--- + +Every version of the [enterprise appliance](/developer-guide/self-hosting/enterprise-appliance) +we have released, newest first. Developer → Self Host shows which one your team recorded as +running and links here when a newer one is available. + + + Versions are immutable. A tag always refers to the same image — we never + re-point one at different content, so a deployment pinned to a tag keeps + getting the bytes it was tested with. Fixes ship as a new version. + + +## How to read a version + +``` +s2.1-pro-20260803-offline +└──┬───┘ └──┬───┘ └──┬──┘ + │ │ └─ variant: `offline` records usage to a local signed ledger + │ └─ the date we published it + └─ the model generation it serves +``` + +--- + +## s2.1-pro-20260803-offline + + +First generally available build of the appliance. + +- Serves the S2.1 Pro voice model, weights baked in — no network access at runtime. +- Two GPUs: language model on GPU 0, vocoder on GPU 1. +- Usage recorded to a local signed ledger; nothing is reported off the host. +- Runs fully non-root (UID 1000). + +**Upgrading:** nothing to migrate — this is the first release. + + + +--- + +## Getting an older version + +The dashboard lists the versions we currently recommend. If you need one that is no longer +listed — to reproduce a node exactly as it was, for instance — contact your account manager +rather than guessing a tag. diff --git a/docs.json b/docs.json index 209091a..eab9bd9 100644 --- a/docs.json +++ b/docs.json @@ -203,7 +203,9 @@ "pages": [ "developer-guide/self-hosting/local-setup", "developer-guide/self-hosting/docker-deployment", - "developer-guide/self-hosting/running-inference" + "developer-guide/self-hosting/running-inference", + "developer-guide/self-hosting/enterprise-appliance", + "developer-guide/self-hosting/enterprise-releases" ] }, {