Skip to content

Folders and files

NameName
Last commit message
Last commit date

Latest commit

 

History

15 Commits
 
 
 
 
 
 

Repository files navigation

AI Python RAG Setup

This project was created as part of our participation in the 18th Student Conference (2026), which took place at Noesis and was organized by the 1st EK of Evosmos. This work was completed by the students Lefteris Trompakas and Asteris Tsiboukas as Kapelo Team, under the supervision of teachers Zoe Belli and George Arnaoutoglou.

What is RAG?
Retrieval-Augmented Generation (or RAG), which loads a specific dataset into an AI model. Without it, a highly intricate and complex model training process would be required. Essentially, RAG converts text into vectors and then splits it into chunks, enabling the model to understand what that information represents. Also you can search more informations in Web

Setup Ollama

  1. Go to ollama.com and select a model
  2. Execute command ollama run [full model name] --> Example ollama run ilsp/llama-krikri-8b-instruct:q4_k_m

Setup RAG File - Linux (Python 3)

  1. cd [directory of use]
  2. python -m venv venv
  3. ./venv/bin/pip install --upgrade pip --only at install (one time)
  4. ./venv/bin/pip install chromadb pypdf python-pptx tqdm requests pytesseract pdf2image pillow mwparserfromhell --only at install (one time)
  5. ./venv/bin/python Setup.py (Setup file must be in the same folder with cd command)

Parameters

--model [ollama model name*] --> Example ./venv/bin/python Setup.py --model krikri-gpu:latest
After load of model Setup.py support indexing from a database (.xml.bz2 dump), like Wikipedia (Greek Wikipedia). Example:
💬 Question: wiki
📂 Path from .xml.bz2 dump: [full file path of .xml.bz2]
*Ollama model name can be found be the execute of command ollama list

Abilities

  • Support OCR from images, pdfs, pptx
  • In the index, a selection is made between sources containing more theory and sources that primarily feature logical content (Mathematics - Computer Science). 'OneTime' refers to the folders containing sources, meaning data that will be added to the model without indexing
  • It supports the automatic extraction of files to the output folder
  • Multithread
  • Progress check
  • 100% locally executed (No API keys required)

Paths

Full file path folder

|-- AuthorizedMath/      (Sources with Math/Logic for indexing)
|-- AuthorizedThe/       (Sources with Theory for indexing)
|-- db/                  (ChromaDB database directory)
|-- materials/           (Materials/Documents storage)
|-- OneTimeMath/         (One-time-use Math sources - no indexing)
|-- OneTimeThe/          (One-time-use Theory sources - no indexing)
|-- Output/              (Automatic file extraction folder)
|-- venv/                (Python Virtual Environment)
|-- Modelfile            (Ollama configuration file)
|-- Modelfile.save       (Backup configuration file)
`-- Setup.py             (Main application script)

Credits

All this code was written by Claude and Gemini, while the idea and review were done by Kapelo Team, which consists of Lefteris Trompakas and Asteris Tsiboukas.
Python libraries: chromadb, pypdf, python-pptx, tqdm, requests, pytesseract, pdf2image, pillow, mwparserfromhell

About

Complete Python RAG setup file for AI model training. Ollama supported

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Used by

Contributors

Languages