Skip to content

Repository files navigation

Big Data Engineering Project

End to End Big data engineering project with streamed and batch data analysis using Flume, Kafka, Sqoop, Spark Streaming, Spark SQL, Hive and HBase in Scala with sample datasets.

  1. Load batch dimension tables data text files to MySql
  2. Load history tables data to MySql using Sqoop
  3. Stream JSON file to Kafka using Flume
  4. Spark Streaming to produce real time KPI's using streamed and batch data lookup.
  5. Spark Streaming to HBase to save streamed data
  6. Join 2 streams in Spark Streaming and push to Kafka Topic.
  7. Batch data anayisis - create fact and pivot table by combining batch and streamed data from HBase to Hive.

Process Flow Diagram

About

End to End Big data engineering project with streamed and batch data analysis using Flume, Kafka, Sqoop, Spark Streaming, Spark SQL, Hive and HBase in Scala with sample datasets.

Resources

Stars

0 stars

Watchers

1 watching

Forks

Releases

Packages

Contributors

Languages