Skip to content

Releases: MatheoPinget-dev/BigData

Release list

Release v1.0.0

Choose a tag to compare

@MatheoPinget-dev MatheoPinget-dev released this 29 Oct 19:50

Release v1.0.0

This is the initial stable release of the CHU Big Data Platform.

This version establishes a complete, end-to-end data platform for ingesting, processing, and analyzing healthcare data. It provides a solid foundation for the CHU
(Cloud Healthcare Unit) group to build its data warehousing and business intelligence capabilities.

Key Features

  • End-to-End ETL Pipeline: A robust ETL pipeline orchestrated by Apache Airflow that processes data from raw sources to analysis-ready tables.
    • Bronze Layer: Ingests raw data from PostgreSQL and CSV files into a MinIO-based data lake.
    • Silver Layer: Cleans, transforms, and normalizes the data to ensure quality and consistency.
    • Gold Layer: Aggregates the cleaned data into a star schema, creating fact and dimension tables optimized for analytics.
  • Modern Data Stack: The entire platform is containerized using Docker and Docker Compose, ensuring portability and ease of deployment.
    • Data Processing: Apache Spark is used for large-scale, distributed data processing.
    • Orchestration: Apache Airflow manages the complex workflow of ETL jobs.
    • Storage: MinIO serves as a scalable S3-compatible data lake.
    • BI & Visualization: Apache Superset is integrated for creating interactive dashboards and charts.
    • Querying: Trino provides a high-performance, distributed SQL query engine to directly query the data lake.
  • Scalable and Secure: The architecture is designed with scalability in mind and includes basic security measures for accessing services.
  • Comprehensive Documentation: The project includes a detailed README.md in both English and French.

How to Use

The platform is launched via a single docker-compose up -d command. Once running, the following services are available: