Skip to content

v0.1.0-alpha.0 - Green Infinuum

Choose a tag to compare

@jdevcs jdevcs released this 28 Aug 10:43
· 21 commits to main since this release
345c79c

RAG Loom - RAG Platform Kit v0.1.0-alpha.0 - Green Infinuum – Release Notes

  • Overview

    • FastAPI microservice implementing Retrieval-Augmented Generation (RAG): document ingestion → chunking → embedding → vector store → semantic search → LLM generation with retrieved context.
  • Key Features

    • Ingestion: Upload PDF/TXT, clean, chunk (sliding window), batch-embed, and store in vector DB.
    • Semantic Search: Vector similarity search with configurable top_k and similarity threshold (with fallback).
    • RAG Generation: Retrieves top-K relevant chunks and calls an LLM with grounded context; returns answer + sources.
    • Pluggable Providers:
      • Embeddings: Sentence-Transformers (local, default), OpenAI (text-embedding-), Cohere (embed-english-).
      • Vector Stores: ChromaDB (local), Qdrant (HTTP/gRPC), Redis (vector index).
      • LLMs: Ollama (default), OpenAI, Cohere, Hugging Face Transformers.
    • Black-box E2E Tests: Real HTTP tests against infra (Qdrant, Ollama) started via docker compose.
    • Unit and Integration tests for functions used in endpoints
    • CI/CD: GitHub Actions workflow to run the e2e suite; caches Docker images to speed up runs.
  • API Endpoints

    • GET /health: Service and platform metadata (vector store, embedding model, llm provider).
    • POST /api/v1/ingest: Multipart upload of PDF/TXT → returns chunks_created and document_id.
    • POST /api/v1/search: Body: { query, top_k, similarity_threshold, filters? } → list of relevant chunks.
    • POST /api/v1/generate: Body: { query, search_params?, context?, temperature, max_tokens } → answer + sources.
  • Configuration

    • Env-driven via app/core/config.py and .env:
      • VECTOR_STORE_TYPE: chroma | qdrant | redis
      • CHROMA_PERSIST_DIRECTORY
      • QDRANT_URL, QDRANT_API_KEY
      • REDIS_URL
      • EMBEDDING_MODEL, EMBEDDING_DIM (default 384)
      • LLM_PROVIDER: ollama | openai | cohere | huggingface
      • OLLAMA_BASE_URL, OLLAMA_MODEL
      • TOP_K, SIMILARITY_THRESHOLD
      • SERVICE_HOST, SERVICE_PORT
    • Defaults: Qdrant at http://localhost:6333, Ollama at http://localhost:11434, local sentence-transformers.
  • Requirements

    • Python 3.12
    • Docker + Docker Compose (for infra/e2e)
    • For OpenAI/Cohere/HF providers, set the relevant API keys in .env.
  • Quick Start (Service)

    • Local venv (3.12) + run:
      • python3.12 -m venv venvpy312 && source venvpy312/bin/activate
      • pip install -r requirements.txt
      • uvicorn app.main:app --host 0.0.0.0 --port 8000
    • Or use helper:
      • bash utilscripts/quick_start.sh start
      • bash utilscripts/quick_start.sh stop
  • Quick Start (Infra + E2E)

    • Start infra + run tests (black-box):
      • bash utilscripts/e2e_run.sh
    • What it does:
      • docker compose up qdrant, redis, ollama (under tests/)
      • waits on health (Qdrant, Ollama), ensures a small Ollama model is pulled
      • sets up Python 3.12 venv, installs tests/requirements-test.txt
      • starts the API locally, waits for /health
      • runs pytest -m e2e
      • stops API and stops containers (keeps them cached for faster re-runs)
  • CI/CD

    • .github/workflows/e2e.yml:
      • Caches Docker images (Qdrant, Redis, Ollama) between runs.
      • Executes utilscripts/e2e_run.sh.
      • On failure, uploads docker compose logs artifact.
      • Tears down infra when finished.
    • .github/workflows/test.yml:
    • for unit and integration testing
  • Notes and Limitations

    • Default small Ollama model used for CI speed; switch OLLAMA_MODEL for higher quality.
    • For OpenAI/Cohere/HF, provide API keys; otherwise the local embedding/LLM options are used.
    • Retrieval applies threshold; if no result passes, it falls back to top results to avoid empty context.

Project Summary
Documentation
Tests