Skip to content

AI Models

Lef edited this page Oct 4, 2026 · 2 revisions

AI Models

Everything runs locally by default. The models live on the host (LM Studio and Ollama), and the backend reaches them over localhost because it uses the host network. Cloud providers are optional and only work with API keys an admin enters.

Roles

The backend doesn't call "a model", it calls a role. Each role has a provider and a model, set in Admin → AI:

Role Default Used for
Storyteller (English) LM Studio, whatever model is loaded (google/gemma-4-e2b) the AI Storyteller in English
Storyteller (Greek) LM Studio llama-krikri-8b-instruct the AI Storyteller in Greek
Utility Ollama llama3.2:3b slash commands, summaries, location suggestions
Classifier Laya if installed, otherwise the loaded LM Studio model OOC moderation, message intent

If a role's provider fails (down, timeout, bad key), the backend falls back to the local storyteller, all inside a 55 second budget so the page never hangs.

Language

The Storyteller answers in the player's language:

  • Greek letters in the message → Greek, and the Greek role (Krikri) answers.
  • Latin text that isn't clearly English (for example Greeklish) follows the player's interface language.
  • Krikri is the only installed model that writes clean Greek. Gemma and Qwen mix in characters from other scripts. If a Greek reply still contains them, the backend retries once and then cleans them out.

VRAM: Krikri needs about 5 GB. If something else holds the GPU (ComfyUI uses about 7 GB here) and Krikri can't load, Greek messages fall back to the English model for 5 minutes. With a busy GPU the simplest setup is to load only Krikri in LM Studio: the English role follows whatever model is loaded, so Krikri then answers both languages.

Memory and rule books (embeddings)

All memory and rule-book search uses text-embedding-bge-m3 in LM Studio. It's multilingual, so a Greek question finds an English passage and the other way round. nomic-embed-text-v1.5 was tested and can't tell Greek texts apart.

If you change the embedding model (EMBEDDING_MODEL in .env), the backend re-embeds the stored collections at startup. Admins can also do it by hand from Admin → AI → Embeddings.

The Storyteller prompt

Every reply gets: the chronicle's rules summary (Classic or V5), your character sheet, the room, the recent conversation, and relevant older scenes from memory. It's fitted to the model's context (8k for the default models). Older lines are trimmed first. If the model still says it's too long, the same model is retried with less history rather than dropping to a smaller model.

Moderation (OOC rooms)

Messages in an out-of-character room are checked for in-character roleplay:

  • 3 warnings in 7 days → a temporary ban in that chronicle only.
  • Admins, helpers and the chronicle's owner are never checked.
  • The check takes about 0.15 s with the LLM classifier.
  • If no classifier is available, nothing gets flagged (it fails open, so nobody is banned by an outage).

Laya and Jev

  • Laya is an open-weights text classifier that runs on the CPU. A ShadowRealms model (OOC vs in-character, message intent, English and Greek) is trained from ml/laya/ in the repo. Once it's installed, the classifier role uses it automatically.
  • Jev (Typesafe) is the hosted version. To use it, add a Typesafe API key in Admin → AI → Classifier.

Cloud providers (optional)

In Admin → AI → Cloud keys an admin can add an Anthropic or OpenAI API key and point a role at it.

  • Keys are stored encrypted and only ever shown as •••• plus the last 4 characters.
  • Each provider has a "Test connection" button.
  • Subscriptions (Claude, ChatGPT, Cursor) are not API access; you need an API key from the provider's developer console.

Checking from the chat

As an admin, in any chronicle chat:

  • /ai health: is LM Studio, Ollama and ChromaDB up (no generation)
  • /ai model: configured and active provider
  • /ai ping: tiny generation, shows latency and the raw reply
  • /ai help: everything else

Settings in .env

Variable What it does
LM_STUDIO_REASONING_EFFORT none turns off "thinking" on reasoning models (Gemma 4, Qwen 3.5). Otherwise short replies come back empty
EMBEDDING_MODEL default text-embedding-bge-m3
STORYTELLER_ATTEMPT_TIMEOUT / STORYTELLER_TIME_BUDGET per-model timeout (45 s) and total budget (55 s) for a Storyteller reply
STORYTELLER_CONTEXT_TOKENS cap on the prompt budget (defaults to what LM Studio reports)

More detail: docs/AI_SYSTEMS.md.

Clone this wiki locally