Repository navigation
Ollaya
Classifying tickets, emails and messages usually means sending sensitive text to a cloud model and paying per token. Ollaya is an open-source command-line runtime for decision models, inspired by Ollama, that keeps this work on the user's own hardware. A typed question about text or JSON goes through a fine-tuned model in a single forward pass and returns calibrated probabilities for each option, for example intent, urgency and churn risk for one customer message.
The outcome is private, fast decisions with no per-call fees: the project reports latency under 100 ms on consumer GPUs. It runs on macOS with Apple silicon, Windows 10 and 11, Linux on x86-64 and ARM64, WSL 2 and Docker, using NVIDIA CUDA or Apple MLX when available and a CPU fallback otherwise. The project lists 19 model families with open weights on Hugging Face, and is licensed Apache-2.0.
The accuracy claims (0.773 against 0.749 for competitors) come from the project's own benchmark and have not been independently verified. The project describes itself as beta.
The typed-judgment approach is shared with the Jev-based tools JevUltrafast and JevReview, though Ollaya runs models locally rather than calling a hosted API.
Ollaya sits in Assess. Local calibrated classification is a useful capability for privacy-sensitive workflows, but there is no first-person use and the benchmarks are vendor-reported.