Skip to content

Tech Stack

Lef edited this page Oct 4, 2026 · 1 revision

Tech Stack

Everything ShadowRealms AI runs on, the AI models it uses, and the tools it was built with. Versions are the ones in use at v0.10.0 (2026-10-04). Where they're set: docker-compose.yml, backend/requirements.txt, frontend/package.json, the Dockerfiles.

AI

Models at runtime

Role Model Runs on Notes
Storyteller (English) whatever model LM Studio has loaded (on the maintainer's machine: llama-krikri-8b-instruct) LM Studio, host GPU Set per role in Admin → AI system
Storyteller (Greek) llama-krikri-8b-instruct (Llama-Krikri by ILSP) LM Studio Chosen because it writes clean Greek; Gemma 4 E2B answered Greek prompts in English in tests
Utility (dice, combat helpers) llama3.2:3b Ollama, host
Embeddings (memory, rule-book search) text-embedding-bge-m3 (BAAI bge-m3), 1024 dims LM Studio Multilingual: English, Greek and across the two
Chat classifier (OOC moderation, intent) Laya (laya-shadowrealms-chat), a fine-tuned mmBERT-base (307M) CPU, ONNX Runtime inside the backend ~230 ms per message on 4 threads, no GPU
Optional cloud Anthropic (default claude-sonnet-5-5) and OpenAI (default gpt-6.1-sol) Their APIs, with an admin-set key Off by default; keys stored encrypted
Optional classifier Typesafe Jev (jev-latest) Typesafe API, with a key Alternative to Laya

Fallbacks: cloud → the role's local model → LM Studio's loaded model → Ollama (and back). Details in AI_SYSTEMS.md.

How Laya was made

  • Base model: mmBERT-base (laya-multilingual, MIT), trained with the RecMeets/Laya training loop (Laya words model on Hugging Face).
  • Data: 332 English and 331 Greek hand-written seed messages (60 of them Greeklish), expanded with synthetic messages from local LLMs in LM Studio: English generated by google/gemma-4-e2b, Greek by llama-krikri-8b-instruct. Each generated message was labelled again by both models (a 2-of-3 label vote with the label it was generated for) and deduplicated.
  • Exported to ONNX and run with ONNX Runtime and tokenizers; no torch at runtime.
  • Training pipeline: ml/laya/. Evaluation on real chat: laya/EVALUATION.md.

RAG and memory

  • ChromaDB 1.5.9 (server image and Python client), cosine distance, every collection embedded with bge-m3.
  • Collections for campaign, character, world, session and message memory, and rule books per edition.
  • The Storyteller prompt is token-budgeted: fixed parts first, then RAG, room memory and history.

Backend

Language Python 3.12 (python:3.12-slim image)
Web framework Flask 3.1, Flask-CORS 6, Flask-JWT-Extended 4.7, Flask-Limiter 4
App server gunicorn 26 (2 workers × 48 threads, gthread)
Database PostgreSQL 16 (postgres:16-alpine), psycopg2
Cache, rate limits, lockouts Redis 7 (redis:7-alpine), redis-py 7
Vector store ChromaDB 1.5.9
Security bcrypt 5 (passwords), cryptography 50 (Fernet + HKDF for stored API keys), JWT access tokens (30 min) with HttpOnly refresh cookies
Classifier runtime onnxruntime 1.30, tokenizers 0.23, numpy 2.5
Live updates Server-sent events with one-time tickets
Dice Own engines for Classic (oWoD Revised) and V5, rolled on the server

Frontend

UI React 18
Build Vite 8 (Rolldown, Oxc), Node 22
Routing react-router 7
Animation motion 14 (formerly Framer Motion)
Translations i18next 26 + react-i18next 17, English and Greek
Charts and dice visuals d3-scale, d3-shape
HTML sanitising DOMPurify
Styles Own design system (frontend/src/design: tokens, 97 SVG glyphs, atmosphere effects) plus Tailwind CSS 3
Fonts Cinzel, Alegreya, EB Garamond, Inter, JetBrains Mono (Google Fonts, SIL OFL)
Tests and lint Jest 30, React Testing Library 16, jest-dom 7, ESLint 9 (no warnings allowed)

Infrastructure

Containers Docker Compose: nginx, backend, PostgreSQL, Redis, ChromaDB, monitoring; the frontend container only for development (--profile dev)
Web server nginx (nginx:alpine): serves the built frontend, proxies /api, strict Content-Security-Policy
GPU stats monitoring container (Python 3.11, psutil, GPUtil, nvidia-ml-py) with the NVIDIA container runtime
Public site an nginx reverse proxy with Let's Encrypt TLS on a separate DietPi box in front of the machine running the stack
Host Linux (Arch-based Omarchy), AMD Ryzen 9 5950X, NVIDIA RTX 4080 SUPER 16 GB

CI, security and project tools

  • GitHub Actions: Python syntax and undefined names (ruff), backend unit tests (pytest), PostgreSQL schema + migration check, frontend tests + lint + build.
  • CodeQL (Python and JavaScript), Dependabot (monthly, majors by hand), secret scanning with push protection, branch protection on main.
  • GitHub Wiki and a GitHub Project board for the roadmap.
  • Headless Chromium (over the DevTools protocol) for UI checks and the README screenshots.

How it was built

The code was written with the help of AI coding tools:

In v0.10 every phase was built in its own branch, reviewed twice (the second review by a different model, Claude Fable 5.1), tested in CI, and checked on the running site before it was released. The plans and logs are in ROADMAP_v0.9.md and ROADMAP_v0.10.md.

This page mirrors docs/TECH_STACK.md.

Clone this wiki locally