Web Crawling project

A web crawler and API for scraping book data built with python and its libraries. This project used MongoDb to store the database; specifically MongoDb Atlas so as to have it on the cloud.
The data retrieved is served using FastAPI with API key aunthentication for security purposes and controlled access.

Features

Asynchronous web crawler application
Stores data in MonogDb
FastApi endpoints to:
- List books ( with pagination)
- Search books with certain filers
- Retrieve specific books by their ID
- View changes
- See statistics of data
Swagger UI documentation (/docs)

Folder Structure

Setup Instructions

Clone the repository
Create and activate a virtual environment
Install dependencies
Create a .env file in the root directory this will hold the authentication keys.

Running the Project

Start the FastAPI server
Access the API
Swagger UI: http://127.0.0.1:8000/docs

ReDoc: http://127.0.0.1:8000/redoc
4. Endpoints

Example MongoDB Document:
{ "_id": "64b9e5f9e5a4b1234567890a",
"title": "A Light in the Attic",
"price_including_tax": 51.77,
"availability": "In stock",
"rating": 3,
"url": "https://books.toscrape.com/catalogue/a-light-in-the-attic_1000/index.html",
"category": "Poetry"
}

Testing

All tests are in the Tests/ folder and run using pytest
Test include:

Endpoints status check
Valid and invalid id checks
Search and filter checks
Categories, stats and health checks

Screenshots

Pytest test

Page Crawling (Page 51 does not exist hence the error message and after 3 tries it moves on) Crawler successfully found 1000 books from 50 pages
Each book and their details crawled and stored to the database successfully

Name		Name	Last commit message	Last commit date
Latest commit History 18 Commits
API		API
Crawler		Crawler
Schelduler		Schelduler
Tests		Tests
Utilities		Utilities
.gitignore		.gitignore
README.md		README.md
requirements.txt		requirements.txt

Provide feedback

Saved searches

Use saved searches to filter your results more quickly

Repository files navigation

Web Crawling project

Features

Folder Structure

Setup Instructions

Running the Project

Testing

Screenshots

About

Uh oh!

Releases

Packages

Uh oh!

Contributors

Uh oh!

Languages

Folders and files

Latest commit

History

Repository files navigation

Web Crawling project

Features

Folder Structure

Setup Instructions

Running the Project

Testing

Screenshots

About

Resources

Uh oh!

Stars

Watchers

Forks

Releases

Packages 0

Uh oh!

Contributors

Uh oh!

Languages

Packages