Production-style enterprise document search & RAG platform — multi-tenant workspaces, hybrid keyword + vector retrieval, source-cited answers, and a BullMQ-backed ingestion pipeline.
This project demonstrates senior full-stack and AI-engineering work end to end: authenticated workspaces, document ingestion with background workers, pgvector similarity search fused with full-text keyword search, retrieval-augmented answers with citations, and an operational dashboard.
- Architecture
- Tech stack
- Quick start
- Environment variables
- Ingestion pipeline
- Search architecture
- RAG flow
- Provider chain
- Security & tenant isolation
- Testing
- Known limitations
- Future improvements
flowchart LR
U[User Browser] -->|HTTPS| W[Next.js App Router]
W -->|better-auth session| AUTH[(Auth tables)]
W -->|Prisma| PG[(PostgreSQL + pgvector)]
W -->|BullMQ enqueue| RDS[(Redis)]
WORKER[Ingestion Worker] -->|consume| RDS
WORKER -->|extract → chunk → embed| PROV{Provider chain}
WORKER -->|write chunks + vectors| PG
W -->|search & ask| PG
W -->|chat| PROV
subgraph Providers
PROV -->|primary| CLA[Claude Agent SDK]
PROV -->|fallback| OL[Ollama gemma3 / nomic-embed-text]
PROV -->|final fallback| MK[Deterministic mock]
end
Two long-running processes: the Next.js app (HTTP) and a separate BullMQ worker
(pnpm worker). They share the Postgres database and Redis queue.
| Layer | Choice |
|---|---|
| Framework | Next.js 16 (App Router) · React 19 · TypeScript |
| Styling | Tailwind CSS v4 (hand-rolled UI primitives) |
| Database | PostgreSQL 16 + pgvector (Prisma ORM, vector(768)) |
| Queue | Redis 7 + BullMQ (ingestion worker, retries, retention) |
| Auth | better-auth (email + password) |
| LLM | Claude Agent SDK → Ollama gemma3:4b → mock (fallback) |
| Embeddings | Ollama nomic-embed-text (768-dim) → deterministic mock |
| File parsing | pdf-parse, mammoth (DOCX), native UTF-8 (TXT/MD) |
| Validation | Zod |
cp .env.example .env
docker compose upThe default Compose stack builds the production image, runs app-migrate
(pnpm db:setup), starts the Next.js app on http://localhost:3000, and runs
the BullMQ worker. The app healthcheck is exposed at /api/health.
Demo login: demo@sourcelens.dev / sourcelens-demo
Optional profiles:
# optional: local LLM/embeddings
docker compose --profile ollama up -d ollama
docker exec -it $(docker ps -qf name=ollama) ollama pull nomic-embed-text
docker exec -it $(docker ps -qf name=ollama) ollama pull gemma3:4b
# optional: S3-compatible storage for uploads
docker compose --profile s3 up -d minio minio-create-bucketFor faster iterative development on macOS or Windows, run dependencies in Compose and the app on the host:
docker compose up -d postgres redis
pnpm install
pnpm db:setup # db push + pgvector + index + seed demo data
pnpm dev
pnpm workerdb:setup creates tables, applies the pgvector/FTS indexes, and seeds demo data.
See .env.example for the full list. Notable:
| Var | Purpose |
|---|---|
DATABASE_URL |
Postgres connection (must have pgvector available). |
REDIS_URL |
Redis for BullMQ. |
BETTER_AUTH_SECRET |
Server secret for cookie/session signing. |
STORAGE_BACKEND |
local by default; set s3 for S3/R2/MinIO-compatible object storage. |
STORAGE_DIR |
Local upload root when using STORAGE_BACKEND=local. |
S3_* |
Endpoint, region, bucket and credentials for the S3 backend. |
INTERNAL_ADMIN_EMAILS |
Comma-separated emails allowed into internal operator tools. |
ANTHROPIC_API_KEY |
Optional. Enables Claude Agent SDK chat. |
RERANKER |
none by default; cohere, voyage, or ollama reranks Ask retrieval. |
OTEL_EXPORTER_OTLP_ENDPOINT |
Enables OpenTelemetry trace export when set, e.g. http://localhost:4317. |
OLLAMA_HOST |
Defaults to http://localhost:11434. |
EMBEDDING_DIM |
Must match the vector(N) column in schema.prisma. |
MAX_UPLOAD_BYTES |
Per-file upload size limit. |
upload → POST /api/documents
→ saveUpload (local FS or S3-compatible object storage)
→ Document row {status: uploaded}
→ enqueueIngest(documentId)
worker ← BullMQ "ingest" queue
→ extractText (pdf-parse / mammoth / utf-8)
→ chunkText (~2000 chars, paragraph-aware, 200-char overlap)
→ embedTexts (Ollama nomic-embed-text → mock)
→ INSERT chunks + vectors (single transaction, raw SQL for pgvector)
→ Document {status: indexed, ingestDurationMs}
→ IngestJob {state: completed}
Failures: BullMQ retries 3× with exponential backoff. Final failure flips
Document.status = failed and writes the error message. The Documents and Jobs
pages surface one-click Retry actions, including Retry all failed on the
Jobs page.
Bull Board is mounted at /internal/bull for workspace owners and emails in
INTERNAL_ADMIN_EMAILS. It exposes the raw ingest queue, including delayed,
stalled, failed, and Redis-level queue details. The Jobs table links each BullMQ
job id to its Bull Board detail page.
OpenTelemetry tracing is disabled unless OTEL_EXPORTER_OTLP_ENDPOINT is set.
When enabled, SourceLens emits Next.js auto-instrumentation plus manual spans for
ingest.job, ingest.extract, ingest.chunk, ingest.embed.batch,
ingest.insert, search.keyword, search.vector, search.fuse,
ask.retrieve, ask.llm, and ask.persist. BullMQ jobs carry traceparent
metadata so worker ingestion spans continue the upload request trace.
Local Jaeger:
docker compose --profile observability up -d jaeger
OTEL_EXPORTER_OTLP_ENDPOINT=http://localhost:4317 pnpm devJaeger UI opens at http://localhost:16686.
Three modes on POST /api/search:
- keyword — Postgres
to_tsvector('english', text) @@ websearch_to_tsqueryranked byts_rank. Uses thechunk_text_ftsGIN index. - vector — pgvector cosine distance (
embedding <=> query_vector), HNSW index for ANN. - hybrid (default) — top-25 of each, fused with reciprocal rank fusion (k = 60).
All queries are scoped by workspaceId in the WHERE clause — there is no path
by which a chunk from another workspace can appear in a result.
Uploads go through src/lib/storage, which selects an adapter with
STORAGE_BACKEND.
localstores files underSTORAGE_DIRand recordsDocument.storagePathas a relative filesystem path such asworkspaceId/object-name.pdf.s3streams uploads through the Next.js server to any S3-compatible API and recordsDocument.storagePathass3://bucket/key.azureis reserved for a later adapter and fails fast if selected.
For local MinIO testing:
docker compose --profile s3 up -d minio minio-create-bucketThen set:
STORAGE_BACKEND=s3
S3_ENDPOINT=http://localhost:9000
S3_REGION=us-east-1
S3_BUCKET=sourcelens
S3_ACCESS_KEY_ID=sourcelens
S3_SECRET_ACCESS_KEY=sourcelens-secret
S3_FORCE_PATH_STYLE=truePOST /api/ask (Ask page) does:
- Hybrid search top-6 chunks for the question.
- Compose a system prompt + numbered context block.
- Optionally rerank the fused top-20 chunks (
RERANKER=cohere|voyage|ollama) before taking the final context window. - Call the chat provider chain.
- Persist the (
question,answer, citations, model, provider, retrievalScore, retrievalMs, rerankMs, llmMs) row inQuestion. - Return the answer plus citations and full context text for the UI.
Citations render directly under the answer with score + filename + chunk number, so any sentence can be verified against its source.
Embeddings and chat use ordered fallback chains. The first provider that succeeds wins; on failure or unavailability the next provider is tried.
Embeddings:
- Ollama
nomic-embed-text(768 dim) — used when/api/tagsresponds. - Deterministic SHA-256-derived mock vector — always available.
Chat:
- Claude Agent SDK (
@anthropic-ai/claude-agent-sdk) whenANTHROPIC_API_KEYis set. - Ollama
gemma3:4blocal model when reachable. - Mock — returns the top retrieved chunks as a labelled
[DEMO MODE]answer, so the UI always renders something usable.
Rerankers:
Reranking is off by default for raw /api/search and on for /api/ask.
search() accepts { rerank: true } when callers want the second stage.
| Provider | Env | Tradeoff | Approximate cost shape |
|---|---|---|---|
| None | RERANKER=none |
Zero latency/cost; preserves RRF order. | Free. |
| Cohere | RERANKER=cohere + COHERE_API_KEY |
Strong hosted cross-encoder with simple search-unit billing. | One search unit for a top-20 rerank. Cohere defines one unit as one query over up to 100 ranked documents, with long docs counted as chunks. |
| Voyage | RERANKER=voyage + VOYAGE_API_KEY |
Hosted reranker with large-context models. | Token-based: query tokens are counted once per document plus document tokens; top-20 short chunks are typically a small fraction of 1M tokens. |
| Ollama | RERANKER=ollama |
Local/private if an Ollama-compatible rerank endpoint is available. | No API fee; pays local CPU/GPU time. |
The chain means the project runs end-to-end with no paid API keys and
without an internet connection (set up Ollama with gemma3 and
nomic-embed-text for real models, or accept the deterministic mock for tests).
- Every API route resolves the caller's user via better-auth and the current
workspace via
requireCurrentWorkspace(). There is no path to query another workspace's data. - All SQL — including the raw vector and full-text queries — joins on
workspaceIdso an attacker manipulating a chunk id cannot fetch foreign data. - Uploaded files are size-checked (
MAX_UPLOAD_BYTES) and type-checked against an allowlist of MIMEs / extensions. - Storage paths are workspace-prefixed and stem-sanitised before being written.
- Per-user token-bucket rate limits on upload / search / ask / retry; 429 with
Retry-After+X-RateLimit-*headers. Configurable viaRATE_LIMIT_*env. - Per-IP anonymous limits on sign-in / sign-up / password reset / verification
resend / public invitation lookup.
TRUST_PROXY=1enablesX-Forwarded-For/Forwardedparsing; without it the limiter falls back to a single shared bucket to refuse spoofed identities. - Email verification (opt-in via
BETTER_AUTH_REQUIRE_EMAIL_VERIFICATION=true) and password reset both go through the email provider chain (#22). Reset links expire in 1 hour, verify links in 24 hours. - Prompt-injection guard (
src/lib/rag/sanitise.ts) inspects every retrieved chunk before it reaches the LLM, flags risky patterns (ignore_previous,role_directive,tag_break,boundary_marker,long_base64,hidden_unicode) and either warns, strips or blocks based onRAG_INJECTION_MODE. Each chunk is wrapped in a<source warnings="...">block in the prompt so the model has explicit metadata about untrusted input, and the Ask page surfaces flags as badges next to each citation. - Secrets never reach the client: only
BETTER_AUTH_URLand public Next.js values are exposed.
pnpm test # Vitest unit suite (≈ 60 tests, sub-second)
pnpm test:coverage # with v8 coverage
pnpm test:e2e # Playwright smoke (signup → upload → ingest → search → ask)Unit tests cover the chunker, RRF fusion, provider chain demotion, streaming chain demotion, the rate-limit Lua semantics, the RBAC rank helper, both Zod schemas, the workspace auth helpers, and the deterministic mock embeddings.
The Playwright smoke is a local-only check. It spins up the dev server and
the BullMQ worker via Playwright's webServer array; both processes are torn
down when the run finishes. CI lives at .github/workflows/ci.yml and runs the
unit suite against Postgres + Redis service containers.
- DOCX path depends on
mammoth; complex Word features (tables, images) are dropped to plain text. - Server-mediated uploads still route through the Next.js app before reaching local disk or S3. Browser-to-object-storage presigned uploads are a later milestone.
- Invitation emails are real via Resend or SMTP when configured; the
default
consoleprovider logs the email to stdout for dev. Email verification + password reset wiring through better-auth is#14. - Embedding model swaps require a full re-ingest until #17 lands.
- Workspace switcher and invitations UI.
- Streaming RAG answers (SSE) and inline citation highlighting on hover.
- Re-ranker pass over the fused top-N before sending to the LLM.
- Per-document delete from search index without re-running the worker.