ContextMemory 0.2.0-beta — self-host launch
Self-hosted memory gateway for your llama.cpp / vLLM server. One OpenAI-compatible /v1 URL. Memory you can open as markdown.
Try it (3 commands)
git clone https://github.com/Kortexio/ContextMemory.git && cd ContextMemory
docker compose -f docker-compose.yml -f docker-compose.llamacpp.yml up --build -d
./scripts/aha-chat.shThe second chat request carries only the new question — the answer comes from gateway memory. CI runs the same check against a real llama-server on every push.
Highlights
- Engines:
docker-compose.llamacpp.yml(CPU) anddocker-compose.vllm.yml(GPU), with tool-calling flags documented (--jinja, vLLM parsers) - Chat aha:
scripts/aha-chat.sh/.ps1— two-turn session recall without client-side history - e2e CI:
e2e-llamacppbadge on main - Docs: README leads with self-host; Engines section in
docs/self-host.md - Hygiene: cleaner Issues (no bot cadence posts), Conventional Commit PR titles, issue templates
Changelog (selected)
Features
- Bundle llama.cpp and vLLM compose overrides with an end-to-end memory check
- Self-host launch readiness (README, engines, e2e, public signals)
- Agentic: raise wiki call budget to 5 distinct queries
Bug fixes
- Agentic wiki-budget loops and repeated response sections
- MCP NDJSON framing for azure-monitor-mcp
- Migrations: cast ConfigJson updates to jsonb
- ArgumentOutOfRange on agentic string truncation
Full compare: v0.1.0-beta...v0.2.0-beta
Note: This is still a beta. Session memory works with small instruct models (CI uses 1.5B). The full agentic tool loop needs larger models with native tool calling.