Multi-agent RAG platform combining vector search, knowledge graphs, and LLM orchestration for document-grounded enterprise Q&A.
- What is MemMesh?
- Architecture
- Quick Start
- Usage
- Configuration
- API Reference
- Testing
- Project Structure
- Troubleshooting
- Contributing
- License
MemMesh is a multi-agent RAG (Retrieval-Augmented Generation) engine that answers questions over your documents. Unlike simple RAG pipelines that only do vector search, MemMesh uses a LangGraph-orchestrated agent pipeline that:
- Rewrites queries for better retrieval
- Routes to vector search (Weaviate), knowledge graphs (Neo4j), or both
- Validates retrieved context with Corrective RAG (CRAG)
- Falls back to web search when context is insufficient
- Cites sources with numbered references
- Supports multi-tenant teams with isolated data per team
- Upload PDFs, DOCX, PPTX, images, HTML — layout-aware parsing via Docling
- Ask questions in natural language, get cited answers with source references
- Switch between conversations with branching message trees
- Export chat sessions as Markdown, JSON, or PDF
- Monitor system health, analytics, and RAG evaluation scores from a dashboard
- Not a general-purpose chatbot — optimized for document-grounded Q&A
- Not a vector database replacement — uses Weaviate under the hood
- Not designed for single-user local use — built for multi-tenant deployments
- No support for real-time document collaboration or editing
- No native mobile app (responsive web UI only)
┌──────────────┐ ┌─────────────────────────────────────────────────────────────────────────────┐
│ Frontend │ │ Backend │
│ (React + │ │ │
│ TanStack) │ │ ┌──────────┐ ┌──────────────────────────────────────────────────────┐ │
│ │ │ │ FastAPI │ │ Agent Pipeline │ │
│ ┌────────┐ │ │ │ Routes │ │ │ │
│ │ Chat │──┼─────┼─►│ │──►│ ┌──────┐ ┌─────────┐ ┌──────┐ ┌──────────┐ │ │
│ │ UI │ │ │ │ /query │ │ │Safety│─►│Rewriter │─►│Router│─►│Retriever │ │ │
│ └────────┘ │ │ │ /upload │ │ │Guard │ │(Gemini) │ │(Gem.)│ │ │ │ │
│ ┌────────┐ │ │ │ /auth │ │ └──────┘ └─────────┘ └──────┘ └────┬─────┘ │ │
│ │ Docs │──┼─────┼─►│ /chat │ │ │ │ │
│ │ UI │ │ │ │ /admin │ │ ┌───────────┼──────────┐ │ │
│ └────────┘ │ │ │ /eval │ │ ▼ ▼ │ │ │
│ ┌────────┐ │ │ └──────────┘ │ ┌──────────────┐ ┌────────────┐ ┌────────────┐ │ │ │
│ │ Admin │──┼─────┼─► Celery │ │ Web Search │◄─│ CRAG │◄─│ Weaviate │ │ │ │
│ │ UI │ │ │ │ Workers │ │ (DuckDuckGo) │ │ (relevance │ │ (vector + │ │ │ │
│ └────────┘ │ │ │ │ │ (fallback) │ │ check) │ │ BM25) │ │ │ │
│ ┌────────┐ │ │ │ ┌─────────┐ │ └──────────────┘ └──────┬─────┘ └────────────┘ │ │ │
│ │ Eval │──┼─────┼─► │ │Ingestion│ │ │ │ │ │
│ │ UI │ │ │ │ │Worker │ │ ▼ │ │ │
│ └────────┘ │ │ │ ├─────────┤ │ ┌──────────────────────────────────────────┐ │ │ │
│ ┌────────┐ │ │ │ │Decay │ │ │ Synthesis │ │ │ │
│ │Analytics│─┼─────┼─► │ │Worker │ │ │ (Gemini generates answer with │ │ │ │
│ │ UI │ │ │ │ ├─────────┤ │ │ [1][2] source citations) │ │ │ │
│ └────────┘ │ │ │ │Analytics│ │ └──────────────────────────────────────────┘ │ │ │
│ │ │ │ │Worker │ │ │ │ │ │
│ │ │ │ ├─────────┤ │ ▼ │ │ │
│ │ │ │ │Drift │ │ ┌──────────────────────────────────────────┐ │ │ │
│ │ │ │ │Worker │ │ │ Output Safety │ │ │ │
│ │ │ │ └─────────┘ │ │ (PII scrub + toxic content check on │ │ │ │
│ │ │ └─────────────┘ │ final answer) │ │ │ │
│ │ │ └──────────────────────────────────────────┘ │ │ │
│ │ │ │ │ │ │
│ │ │ ▼ │ │ │
│ │ │ ┌────────────────────────────────────────────────────────────┐ │ │ │
│ │ │ │ Persistence Layer │ │ │ │
│ │ │ │ ┌─────────┐ ┌──────────┐ ┌──────────┐ ┌──────────┐ │ │ │ │
│ │ │ │ │ MySQL │ │ Weaviate │ │ Neo4j │ │ Redis │ │ │ │ │
│ │ │ │ │ (rel. │ │ (vector │ │ (graph │ │ (cache + │ │ │ │ │
│ │ │ │ │ data) │ │ search) │ │ store) │ │ queue) │ │ │ │ │
│ │ │ │ └─────────┘ └──────────┘ └──────────┘ └──────────┘ │ │ │ │
│ │ │ └────────────────────────────────────────────────────────────┘ │ │ │
│ │ └─────────────────────────────────────────────────────────────────────┘ │ │
└──────────────┘ └───────────────────────────────────────────────────────────────────────┘ │
│
┌─────────────────────────────────────────────────────────────────────────────────────────────────┘
│
▼
curl -X POST http://localhost:8000/api/query \
-H "Authorization: Bearer <token>" \
-d '{"query": "What is revenue growth?", "team_id": "..."}'
| Stage | Component | What it does |
|---|---|---|
| Safety | agents/safety.py |
PII scrubbing + toxic content filtering (input and output) |
| Rewrite | agents/rewriter.py |
Gemini generates 2-3 query variants for broader recall |
| Route | agents/router.py |
Gemini classifies query as vector, graph, or hybrid |
| Retrieve | agents/retriever.py |
Weaviate hybrid search (BM25 + vector) and/or Neo4j Cypher queries |
| CRAG | agents/crag.py |
Gemini evaluates relevance score; falls back to web search (DuckDuckGo) if < 0.5 |
| Synthesize | agents/safety.py |
Final answer generated with [1], [2] source citations |
| Memory | agents/memory.py |
Redis-backed conversation history (7-day TTL, 20 msg sliding window) |
| Store | Technology | Purpose |
|---|---|---|
| Vector DB | Weaviate 1.28 | Document chunk embeddings (3072-dim Gemini), hybrid search, multi-tenancy |
| Graph DB | Neo4j 5 | Entity-relationship knowledge graph, per-tenant database isolation |
| RDBMS | MySQL 8.0 | Users, teams, sessions, messages, source docs, audit logs |
| Cache | Redis 7 | Celery broker + result backend, semantic cache, rate limiter, conversation memory |
| Schedule | Task | Description |
|---|---|---|
| Every hour | Memory decay | Reduces importance scores in Weaviate + Neo4j; deletes entities with score ≤ 0 |
| Every hour | Analytics aggregation | Computes query volume, latency, feedback metrics |
| Daily | Data drift detection | Analyzes user feedback for topic shifts |
| Weekly | DeepEval run | Computes faithfulness, hallucination, answer relevancy scores |
- Python 3.12 (with
uvinstalled:curl -LsSf https://astral.sh/uv/install.sh | sh) - Node.js 20+ and npm
- Docker and Docker Compose v2
- A Gemini API key (get one here)
git clone <repo-url>
cd memmesh
cp .env.example .env
# Edit .env — at minimum set GEMINI_API_KEYdocker compose up -dThis starts Weaviate, MySQL 8.0, Neo4j 5, and Redis 7.
# Backend
uv sync --project backend
# Frontend
cd frontend && npm install && cd ..make db-upgradeOpen three terminal tabs:
# Terminal 1 — Backend API
make dev-backend
# Terminal 2 — Frontend dev server
make dev-frontend
# Terminal 3 — Celery worker (for document ingestion, memory decay, etc.)
make dev-workerVisit http://localhost:5173 and log in with:
- Email:
superadmin@memmesh.com - Password:
admin_secret_password_change_me
curl -X POST http://localhost:8000/api/query \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <token>" \
-d '{
"query": "What does the Q3 financial report say about revenue growth?",
"team_id": "team-uuid-here"
}'curl -N -X POST http://localhost:8000/api/query/stream \
-H "Content-Type: application/json" \
-H "Authorization: Bearer <token>" \
-d '{
"query": "Summarize our competitive analysis findings",
"team_id": "team-uuid-here"
}'curl -X POST http://localhost:8000/api/upload \
-H "Authorization: Bearer <token>" \
-F "file=@report.pdf" \
-F "team_id=team-uuid-here"curl http://localhost:8000/api/documents \
-H "Authorization: Bearer <token>" \
-H "X-Team-ID: team-uuid-here"# Markdown
curl http://localhost:8000/api/chat/sessions/<id>/export?format=md \
-H "Authorization: Bearer <token>"
# JSON
curl http://localhost:8000/api/chat/sessions/<id>/export?format=json \
-H "Authorization: Bearer <token>"curl http://localhost:8000/api/healthAll configuration is via environment variables in .env. See .env.example for a template.
| Variable | Default | Description |
|---|---|---|
GEMINI_API_KEY |
— | Google Gemini API key |
GEMINI_MODEL |
gemini-2.5-flash |
Model identifier |
| Variable | Default | Description |
|---|---|---|
MYSQL_HOST |
127.0.0.1 |
Database host |
MYSQL_PORT |
3306 |
Database port |
MYSQL_USER |
app_user |
Database user |
MYSQL_PASSWORD |
user_password_change_me |
Database password |
MYSQL_DATABASE |
memmesh |
Database name |
| Variable | Default | Description |
|---|---|---|
WEAVIATE_HOST |
127.0.0.1 |
Weaviate host |
WEAVIATE_PORT |
8080 |
HTTP API port |
WEAVIATE_GRPC_PORT |
50051 |
gRPC port |
| Variable | Default | Description |
|---|---|---|
NEO4J_URI |
bolt://127.0.0.1:7687 |
Bolt connection URI |
NEO4J_USER |
neo4j |
Database user |
NEO4J_PASSWORD |
neo4j_password_change_me |
Database password |
| Variable | Default | Description |
|---|---|---|
REDIS_URL |
redis://127.0.0.1:6379/0 |
Redis connection string |
| Variable | Default | Description |
|---|---|---|
JWT_SECRET |
super_secret_jwt_key_change_me |
Signing key |
JWT_ALGORITHM |
HS256 |
Signing algorithm |
JWT_EXPIRY_MINUTES |
60 |
Token TTL |
| Variable | Default | Description |
|---|---|---|
RATE_LIMIT_LOGIN |
30/minute |
Login endpoint |
RATE_LIMIT_REGISTER |
3/minute |
Register endpoint |
RATE_LIMIT_QUERY |
30/minute |
Query endpoints |
RATE_LIMIT_GLOBAL |
100/minute |
Global limit |
| Variable | Default | Description |
|---|---|---|
SEMANTIC_CACHE_TTL |
86400 |
Cache TTL in seconds (24h) |
SEMANTIC_CACHE_THRESHOLD |
0.92 |
Cosine similarity threshold |
| Variable | Default | Description |
|---|---|---|
CORS_ORIGINS |
http://localhost:5173 |
Comma-separated allowed origins |
MAX_UPLOAD_SIZE_MB |
50 |
Max document upload size |
SUPERADMIN_EMAIL |
superadmin@memmesh.com |
Seeded on startup |
SUPERADMIN_PASSWORD |
admin_secret_password_change_me |
Seeded on startup |
All endpoints are prefixed with /api.
| Method | Path | Auth | Description |
|---|---|---|---|
| GET | /api/health |
— | System health (MySQL, Redis, Weaviate, Neo4j) |
| POST | /api/auth/login |
Rate-limited | Login, returns JWT |
| POST | /api/auth/register |
Rate-limited | Register new user |
| POST | /api/auth/refresh |
Token | Refresh JWT |
| POST | /api/query |
Token | Blocking Q&A |
| POST | /api/query/stream |
Token | SSE streaming Q&A |
| GET | /api/chat/sessions |
Token | List sessions |
| GET | /api/chat/messages |
Token | Get session messages |
| GET | /api/chat/sessions/{id}/export |
Token | Export session (md/json/pdf) |
| POST | /api/upload |
Team lead+ | Upload document |
| GET | /api/upload/status/{id} |
Token | Upload task status |
| POST | /api/feedback |
Token | Submit feedback (up/down) |
| POST | /api/eval |
Superadmin | Trigger RAG evaluation |
| GET | /api/eval/scores |
Superadmin | Evaluation history |
| GET | /api/admin/teams |
Superadmin | List teams |
| POST | /api/admin/teams |
Superadmin | Create team |
| DELETE | /api/admin/teams/{id} |
Superadmin | Delete team |
| GET | /api/admin/users |
Superadmin | List users |
| POST | /api/admin/users |
Superadmin | Create user |
| GET | /api/documents |
Token | List team documents |
| GET | /api/documents/{id}/pdf |
Token | Serve PDF file |
# Unit tests (SQLite in-memory, mocked LLM)
make test-backend
# With real MySQL (Docker required)
USE_TESTCONTAINERS=1 make test-backendUses MOCK_LLM=true by default. Set MOCK_LLM=false to run against real Gemini.
make test-frontendmake test-e2eE2E tests require the backend, frontend, and infrastructure to be running. MOCK_LLM=true is set automatically in the Playwright config.
make testMemMesh/
├── backend/
│ ├── agents/ # LangGraph agent pipeline
│ │ ├── graph_orchestrator.py # State graph + SSE streaming
│ │ ├── rewriter.py # Query rewriting
│ │ ├── router.py # Route classification
│ │ ├── retriever.py # Vector + graph retrieval
│ │ ├── crag.py # Corrective RAG relevance check
│ │ ├── safety.py # PII + toxic content
│ │ ├── web_search.py # DuckDuckGo fallback
│ │ ├── memory.py # Conversation history
│ │ └── telemetry.py # SSE event types
│ ├── api/routes/ # FastAPI route handlers
│ ├── auth/ # JWT, middleware, password hashing
│ ├── cache/ # Semantic caching (Redis + embeddings)
│ ├── db/ # Database clients (MySQL, Weaviate, Neo4j)
│ ├── ingestion/ # Document parsing, entity extraction, versioning
│ ├── tasks/ # Celery background workers
│ ├── export/ # Chat session exporters (md/json/pdf)
│ ├── middleware/ # Error handlers, logging
│ ├── migrations/ # Alembic migrations
│ ├── tests/ # Pytest test suite
│ ├── config.py # Pydantic Settings
│ ├── main.py # FastAPI app factory
│ ├── models.py # SQLAlchemy ORM models
│ └── rate_limiter.py # slowapi limiter
├── frontend/
│ ├── src/
│ │ ├── components/ # React components (chat, docs, admin, shared, UI)
│ │ ├── hooks/ # useChatStream, etc.
│ │ ├── routes/ # TanStack Router pages
│ │ ├── utils/ # Auth, query, health, upload helpers
│ │ └── types/ # TypeScript type definitions
│ ├── package.json
│ └── vite.config.ts
├── tests/e2e/ # Playwright E2E tests
├── docs/ # Design specs and plans
├── docker-compose.yml # Infrastructure containers
├── Makefile # Dev commands
└── .env.example # Configuration template
Run uv sync --project backend from the project root. Make sure your terminal is not inside the backend/ directory when starting the server.
Set CELERY_WORKER_PREFETCH_MULTIPLIER=1 and worker_max_tasks_per_child=1 in your environment. MemMesh's Celery app already configures fork safety, but your broker settings may override it.
The login endpoint is rate-limited to 30 requests/minute by default. If you're testing programmatically, reuse the JWT token instead of logging in for every request.
The PDF viewer uses pdfjs-dist with credentials from the auth header. Make sure you're running a recent Chrome/Edge/Firefox. Safari's PDF.js support is limited.
Ensure Docker containers are running (docker compose ps). On first startup, Weaviate takes 15-30 seconds to initialize. Wait for the health check to pass.
Contributions are welcome. Open an issue to discuss changes before submitting PRs.
- Report bugs and suggest features via GitHub Issues
- Follow existing code style — the project uses
rufffor Python andprettierfor TypeScript - Ensure tests pass before submitting:
make test
MIT. See LICENSE for details.