-
Notifications
You must be signed in to change notification settings - Fork 5
Home
Deploy Jina AI's embedding, reranker, and reader models inside customer environments that cannot reach the internet.
![]()
For sales/SA/field engineers walking a customer through their first deployment, start with Why Air-Gap, then Quick Start. For developers integrating the API into an application, start with API Reference.
PHASE 1 (network) PHASE 2 (offline)
───────────────── ─────────────────
┌────────────────┐ USB / SCP / disk ┌────────────────┐ port 8080
│ connected host │ ──────────────────► │ air-gapped host│ ──────► app
│ │ .tar.gz │ │
│ jina-on-prem.py │ │ docker load │ OpenAI / Cohere
│ bundle │ │ docker run │ Gemini / Voyage
└────────────────┘ └────────────────┘
│ │
▼ ▼
download weights serve embeddings,
+ deps from HF Hub reranking, readers
docker build zero outbound calls
That's the whole product. The connected machine has internet to fetch model weights and dependencies. Everything is baked into a single Docker image and exported as a .tar.gz. The offline machine only needs Docker.

- 28 models: Jina embeddings (v5, v4, v3, v2), rerankers, ColBERT, CLIP, ReaderLM, VLM. See Model Catalog.
- 4 API schemas simultaneously: OpenAI, Cohere, Google Gemini, Voyage AI - drop-in for any client.
- Multimodal: text + image + audio + video on omni/clip/v4 models.
- GPU and CPU: same model can be packaged either way.
-
Elasticsearch inference service: works as a
service: openaiendpoint out of the box.
| You are... | Start here |
|---|---|
| An SA/sales engineer evaluating jina-on-prem for a customer | Why Air-Gap, then Customer Scenarios |
| Comparing this against Ollama / vLLM / ONNX / hosted API | Comparison vs alternatives |
| A field engineer deploying at a customer site | Quick Start, then Sizing & Hardware |
| A developer integrating the API | API Reference |
| Building a new bundle from scratch | Bundling Guide |
| Rolling out a new model version | Versioning & Updates |
| Hitting an error | Troubleshooting, FAQ |
Most Jina v5/v4/v3 models are CC-BY-NC-4.0: commercial use needs a license. Contact Elastic sales. v2 and v1 models are Apache-2.0 and free for any use. Per-model license is in the Model Catalog.
Repo - Issues - Releases - Commercial license
Documentation source lives in docs/; edit there via PR and run ./scripts/sync-wiki.sh to mirror to the wiki.