Repository navigation
Release v1.0.0
This is the initial stable release of the CHU Big Data Platform.
This version establishes a complete, end-to-end data platform for ingesting, processing, and analyzing healthcare data. It provides a solid foundation for the CHU
(Cloud Healthcare Unit) group to build its data warehousing and business intelligence capabilities.
Key Features
- End-to-End ETL Pipeline: A robust ETL pipeline orchestrated by Apache Airflow that processes data from raw sources to analysis-ready tables.
- Bronze Layer: Ingests raw data from PostgreSQL and CSV files into a MinIO-based data lake.
- Silver Layer: Cleans, transforms, and normalizes the data to ensure quality and consistency.
- Gold Layer: Aggregates the cleaned data into a star schema, creating fact and dimension tables optimized for analytics.
- Modern Data Stack: The entire platform is containerized using Docker and Docker Compose, ensuring portability and ease of deployment.
- Data Processing: Apache Spark is used for large-scale, distributed data processing.
- Orchestration: Apache Airflow manages the complex workflow of ETL jobs.
- Storage: MinIO serves as a scalable S3-compatible data lake.
- BI & Visualization: Apache Superset is integrated for creating interactive dashboards and charts.
- Querying: Trino provides a high-performance, distributed SQL query engine to directly query the data lake.
- Scalable and Secure: The architecture is designed with scalability in mind and includes basic security measures for accessing services.
- Comprehensive Documentation: The project includes a detailed README.md in both English and French.
How to Use
The platform is launched via a single docker-compose up -d command. Once running, the following services are available:
- Airflow UI: http://localhost:8080
- Jupyter Lab: http://localhost:8888
- MinIO Console: http://localhost:9001
- Superset UI: http://localhost:8088
- Trino UI: http://localhost:8090