v0.1.0-alpha.0 - Green Infinuum
RAG Loom - RAG Platform Kit v0.1.0-alpha.0 - Green Infinuum – Release Notes
-
Overview
- FastAPI microservice implementing Retrieval-Augmented Generation (RAG): document ingestion → chunking → embedding → vector store → semantic search → LLM generation with retrieved context.
-
Key Features
- Ingestion: Upload PDF/TXT, clean, chunk (sliding window), batch-embed, and store in vector DB.
- Semantic Search: Vector similarity search with configurable top_k and similarity threshold (with fallback).
- RAG Generation: Retrieves top-K relevant chunks and calls an LLM with grounded context; returns answer + sources.
- Pluggable Providers:
- Embeddings: Sentence-Transformers (local, default), OpenAI (text-embedding-), Cohere (embed-english-).
- Vector Stores: ChromaDB (local), Qdrant (HTTP/gRPC), Redis (vector index).
- LLMs: Ollama (default), OpenAI, Cohere, Hugging Face Transformers.
- Black-box E2E Tests: Real HTTP tests against infra (Qdrant, Ollama) started via docker compose.
- Unit and Integration tests for functions used in endpoints
- CI/CD: GitHub Actions workflow to run the e2e suite; caches Docker images to speed up runs.
-
API Endpoints
- GET /health: Service and platform metadata (vector store, embedding model, llm provider).
- POST /api/v1/ingest: Multipart upload of PDF/TXT → returns chunks_created and document_id.
- POST /api/v1/search: Body: { query, top_k, similarity_threshold, filters? } → list of relevant chunks.
- POST /api/v1/generate: Body: { query, search_params?, context?, temperature, max_tokens } → answer + sources.
-
Configuration
- Env-driven via
app/core/config.pyand.env:- VECTOR_STORE_TYPE: chroma | qdrant | redis
- CHROMA_PERSIST_DIRECTORY
- QDRANT_URL, QDRANT_API_KEY
- REDIS_URL
- EMBEDDING_MODEL, EMBEDDING_DIM (default 384)
- LLM_PROVIDER: ollama | openai | cohere | huggingface
- OLLAMA_BASE_URL, OLLAMA_MODEL
- TOP_K, SIMILARITY_THRESHOLD
- SERVICE_HOST, SERVICE_PORT
- Defaults: Qdrant at
http://localhost:6333, Ollama athttp://localhost:11434, local sentence-transformers.
- Env-driven via
-
Requirements
- Python 3.12
- Docker + Docker Compose (for infra/e2e)
- For OpenAI/Cohere/HF providers, set the relevant API keys in
.env.
-
Quick Start (Service)
- Local venv (3.12) + run:
- python3.12 -m venv venvpy312 && source venvpy312/bin/activate
- pip install -r requirements.txt
- uvicorn app.main:app --host 0.0.0.0 --port 8000
- Or use helper:
- bash utilscripts/quick_start.sh start
- bash utilscripts/quick_start.sh stop
- Local venv (3.12) + run:
-
Quick Start (Infra + E2E)
- Start infra + run tests (black-box):
- bash utilscripts/e2e_run.sh
- What it does:
- docker compose up qdrant, redis, ollama (under
tests/) - waits on health (Qdrant, Ollama), ensures a small Ollama model is pulled
- sets up Python 3.12 venv, installs
tests/requirements-test.txt - starts the API locally, waits for /health
- runs pytest
-m e2e - stops API and stops containers (keeps them cached for faster re-runs)
- docker compose up qdrant, redis, ollama (under
- Start infra + run tests (black-box):
-
CI/CD
.github/workflows/e2e.yml:- Caches Docker images (Qdrant, Redis, Ollama) between runs.
- Executes
utilscripts/e2e_run.sh. - On failure, uploads docker compose logs artifact.
- Tears down infra when finished.
.github/workflows/test.yml:- for unit and integration testing
-
Notes and Limitations
- Default small Ollama model used for CI speed; switch
OLLAMA_MODELfor higher quality. - For OpenAI/Cohere/HF, provide API keys; otherwise the local embedding/LLM options are used.
- Retrieval applies threshold; if no result passes, it falls back to top results to avoid empty context.
- Default small Ollama model used for CI speed; switch